{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/data-driven-tree-transforms-and-metrics","title":"Data-Driven Tree Transforms and Metrics","arxiv_id":"1708.05768","date":"2017-08-18","proceeding":null,"authors":["Gal Mishne","Ronen Talmon","Israel Cohen","Ronald R. Coifman","Yuval Kluger"],"abstract":"We consider the analysis of high dimensional data given in the form of a\nmatrix with columns consisting of observations and rows consisting of features.\nOften the data is such that the observations do not reside on a regular grid,\nand the given order of the features is arbitrary and does not convey a notion\nof locality. Therefore, traditional transforms and metrics cannot be used for\ndata organization and analysis. In this paper, our goal is to organize the data\nby defining an appropriate representation and metric such that they respect the\nsmoothness and structure underlying the data. We also aim to generalize the\njoint clustering of observations and features in the case the data does not\nfall into clear disjoint groups. For this purpose, we propose multiscale\ndata-driven transforms and metrics based on trees. Their construction is\nimplemented in an iterative refinement procedure that exploits the\nco-dependencies between features and observations. Beyond the organization of a\nsingle dataset, our approach enables us to transfer the organization learned\nfrom one dataset to another and to integrate several datasets together. We\npresent an application to breast cancer gene expression analysis: learning\nmetrics on the genes to cluster the tumor samples into cancer sub-types and\nvalidating the joint organization of both the genes and the samples. We\ndemonstrate that using our approach to combine information from multiple gene\nexpression cohorts, acquired by different profiling technologies, improves the\nclustering of tumor samples.","url_abs":"http://arxiv.org/abs/1708.05768v1","url_pdf":"http://arxiv.org/pdf/1708.05768v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"data-driven-tree-transforms-and-metrics","repo_url":"https://github.com/gmishne/pyquest","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"clustering","task_name":"Clustering"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}