{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/predicting-b-cell-receptor-substitution","title":"Predicting B Cell Receptor Substitution Profiles Using Public Repertoire Data","arxiv_id":"1802.06406","date":"2018-02-18","proceeding":null,"authors":[],"abstract":"B cells develop high affinity receptors during the course of affinity\nmaturation, a cyclic process of mutation and selection. At the end of affinity\nmaturation, a number of cells sharing the same ancestor (i.e. in the same\n\"clonal family\") are released from the germinal center, their amino acid\nfrequency profile reflects the allowed and disallowed substitutions at each\nposition. These clonal-family-specific frequency profiles, called \"substitution\nprofiles\", are useful for studying the course of affinity maturation as well as\nfor antibody engineering purposes. However, most often only a single sequence\nis recovered from each clonal family in a sequencing experiment, making it\nimpossible to construct a clonal-family-specific substitution profile. Given\nthe public release of many high-quality large B cell receptor datasets, one may\nask whether it is possible to use such data in a prediction model for\nclonal-family-specific substitution profiles. In this paper, we present the\nmethod \"Substitution Profiles Using Related Families\" (SPURF), a penalized\ntensor regression framework that integrates information from a rich assemblage\nof datasets to predict the clonal-family-specific substitution profile for any\nsingle input sequence. Using this framework, we show that substitution profiles\nfrom similar clonal families can be leveraged together with simulated\nsubstitution profiles and germline gene sequence information to improve\nprediction. We fit this model on a large public dataset and validate the\nrobustness of our approach on an external dataset. Furthermore, we provide a\ncommand-line tool in an open-source software package\n(https://github.com/krdav/SPURF) implementing these ideas and providing easy\nprediction using our pre-fit models.","url_abs":"http://arxiv.org/abs/1802.06406v1","url_pdf":"http://arxiv.org/pdf/1802.06406v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"predicting-b-cell-receptor-substitution","repo_url":"https://github.com/krdav/SPURF","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}