{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/are-clusterings-of-multiple-data-views","title":"Are Clusterings of Multiple Data Views Independent?","arxiv_id":"1901.03905","date":"2019-01-12","proceeding":null,"authors":["Lucy L. Gao","Jacob Bien","Daniela Witten"],"abstract":"In the Pioneer 100 (P100) Wellness Project (Price and others, 2017), multiple\ntypes of data are collected on a single set of healthy participants at multiple\ntimepoints in order to characterize and optimize wellness. One way to do this\nis to identify clusters, or subgroups, among the participants, and then to\ntailor personalized health recommendations to each subgroup. It is tempting to\ncluster the participants using all of the data types and timepoints, in order\nto fully exploit the available information. However, clustering the\nparticipants based on multiple data views implicitly assumes that a single\nunderlying clustering of the participants is shared across all data views. If\nthis assumption does not hold, then clustering the participants using multiple\ndata views may lead to spurious results. In this paper, we seek to evaluate the\nassumption that there is some underlying relationship among the clusterings\nfrom the different data views, by asking the question: are the clusters within\neach data view dependent or independent? We develop a new test for answering\nthis question, which we then apply to clinical, proteomic, and metabolomic\ndata, across two distinct timepoints, from the P100 study. We find that while\nthe subgroups of the participants defined with respect to any single data type\nseem to be dependent across time, the clustering among the participants based\non one data type (e.g. proteomic data) appears not to be associated with the\nclustering based on another data type (e.g. clinical data).","url_abs":"http://arxiv.org/abs/1901.03905v1","url_pdf":"http://arxiv.org/pdf/1901.03905v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"are-clusterings-of-multiple-data-views","repo_url":"https://github.com/lucylgao/independent-clusterings-code","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null},{"paper_slug":"are-clusterings-of-multiple-data-views","repo_url":"https://github.com/lucylgao/multiviewtest","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"clustering","task_name":"Clustering"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}