{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/fast-two-sample-testing-with-analytic","title":"Fast Two-Sample Testing with Analytic Representations of Probability Measures","arxiv_id":"1506.04725","date":"2015-06-15","proceeding":"NeurIPS 2015 12","authors":["Kacper Chwialkowski","Aaditya Ramdas","Dino Sejdinovic","Arthur Gretton"],"abstract":"We propose a class of nonparametric two-sample tests with a cost linear in\nthe sample size. Two tests are given, both based on an ensemble of distances\nbetween analytic functions representing each of the distributions. The first\ntest uses smoothed empirical characteristic functions to represent the\ndistributions, the second uses distribution embeddings in a reproducing kernel\nHilbert space. Analyticity implies that differences in the distributions may be\ndetected almost surely at a finite number of randomly chosen\nlocations/frequencies. The new tests are consistent against a larger class of\nalternatives than the previous linear-time tests based on the (non-smoothed)\nempirical characteristic functions, while being much faster than the current\nstate-of-the-art quadratic-time kernel-based or energy distance-based tests.\nExperiments on artificial benchmarks and on challenging real-world testing\nproblems demonstrate that our tests give a better power/time tradeoff than\ncompeting approaches, and in some cases, better outright power than even the\nmost expensive quadratic-time tests. This performance advantage is retained\neven in high dimensions, and in cases where the difference in distributions is\nnot observable with low order statistics.","url_abs":"http://arxiv.org/abs/1506.04725v1","url_pdf":"http://arxiv.org/pdf/1506.04725v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"fast-two-sample-testing-with-analytic","repo_url":"https://github.com/kacperChwialkowski/analyticMeanEmbeddings","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"hypothesis-testing","task_name":"Two-sample testing"},{"task_slug":"two","task_name":"Vocal Bursts Valence Prediction"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1506.04725","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}