{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/likelihood-based-inference-of-b-cell-clonal","title":"Likelihood-based inference of B-cell clonal families","arxiv_id":"1603.08127","date":"2016-06-16","proceeding":null,"authors":[],"abstract":"The human immune system depends on a highly diverse collection of\nantibody-making B cells. B cell receptor sequence diversity is generated by a\nrandom recombination process called \"rearrangement\" forming progenitor B cells,\nthen a Darwinian process of lineage diversification and selection called\n\"affinity maturation.\" The resulting receptors can be sequenced in high\nthroughput for research and diagnostics. Such a collection of sequences\ncontains a mixture of various lineages, each of which may be quite numerous, or\nmay consist of only a single member. As a step to understanding the process and\nresult of this diversification, one may wish to reconstruct lineage membership,\ni.e. to cluster sampled sequences according to which came from the same\nrearrangement events. We call this clustering problem \"clonal family\ninference.\" In this paper we describe and validate a likelihood-based framework\nfor clonal family inference based on a multi-hidden Markov Model (multi-HMM)\nframework for B cell receptor sequences. We describe an agglomerative algorithm\nto find a maximum likelihood clustering, two approximate algorithms with\nvarious trade-offs of speed versus accuracy, and a third, fast algorithm for\nfinding specific lineages. We show that under simulation these algorithms\ngreatly improve upon existing clonal family inference methods, and that they\nalso give significantly different clusters than previous methods when applied\nto two real data sets.","url_abs":"http://arxiv.org/abs/1603.08127v2","url_pdf":"http://arxiv.org/pdf/1603.08127v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"likelihood-based-inference-of-b-cell-clonal","repo_url":"https://github.com/psathyrella/partis","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"clustering","task_name":"Clustering"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}