{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/the-shape-of-data-and-probability-measures","title":"The Shape of Data and Probability Measures","arxiv_id":"1509.04632","date":"2015-09-15","proceeding":null,"authors":["Diego Hernán Díaz Martínez","Facundo Mémoli","Washington Mio"],"abstract":"We introduce the notion of multiscale covariance tensor fields (CTF)\nassociated with Euclidean random variables as a gateway to the shape of their\ndistributions. Multiscale CTFs quantify variation of the data about every point\nin the data landscape at all spatial scales, unlike the usual covariance tensor\nthat only quantifies global variation about the mean. Empirical forms of\nlocalized covariance previously have been used in data analysis and\nvisualization, but we develop a framework for the systematic treatment of\ntheoretical questions and computational models based on localized covariance.\nWe prove strong stability theorems with respect to the Wasserstein distance\nbetween probability measures, obtain consistency results, as well as estimates\nfor the rate of convergence of empirical CTFs. These results ensure that CTFs\nare robust to sampling, noise and outliers. We provide numerous illustrations\nof how CTFs let us extract shape from data and also apply CTFs to manifold\nclustering, the problem of categorizing data points according to their noisy\nmembership in a collection of possibly intersecting, smooth submanifolds of\nEuclidean space. We prove that the proposed manifold clustering method is\nstable and carry out several experiments to validate the method.","url_abs":"http://arxiv.org/abs/1509.04632v2","url_pdf":"http://arxiv.org/pdf/1509.04632v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"the-shape-of-data-and-probability-measures","repo_url":"https://bitbucket.org/diegodiaz-math/ctf-files","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"clustering","task_name":"Clustering"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}