{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/parallel-clustering-of-single-cell","title":"Parallel Clustering of Single Cell Transcriptomic Data with Split-Merge Sampling on Dirichlet Process Mixtures","arxiv_id":"1812.10048","date":"2018-12-25","proceeding":null,"authors":["Tiehang Duan","José P. Pinto","Xiaohui Xie"],"abstract":"Motivation: With the development of droplet based systems, massive single\ncell transcriptome data has become available, which enables analysis of\ncellular and molecular processes at single cell resolution and is instrumental\nto understanding many biological processes. While state-of-the-art clustering\nmethods have been applied to the data, they face challenges in the following\naspects: (1) the clustering quality still needs to be improved; (2) most models\nneed prior knowledge on number of clusters, which is not always available; (3)\nthere is a demand for faster computational speed. Results: We propose to tackle\nthese challenges with Parallel Split Merge Sampling on Dirichlet Process\nMixture Model (the Para-DPMM model). Unlike classic DPMM methods that perform\nsampling on each single data point, the split merge mechanism samples on the\ncluster level, which significantly improves convergence and optimality of the\nresult. The model is highly parallelized and can utilize the computing power of\nhigh performance computing (HPC) clusters, enabling massive clustering on huge\ndatasets. Experiment results show the model outperforms current widely used\nmodels in both clustering quality and computational speed. Availability: Source\ncode is publicly available on\nhttps://github.com/tiehangd/Para_DPMM/tree/master/Para_DPMM_package","url_abs":"http://arxiv.org/abs/1812.10048v1","url_pdf":"http://arxiv.org/pdf/1812.10048v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"parallel-clustering-of-single-cell","repo_url":"https://github.com/tiehangd/Para_DPMM","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"clustering","task_name":"Clustering"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}