{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/revisiting-classifier-two-sample-tests","title":"Revisiting Classifier Two-Sample Tests","arxiv_id":"1610.06545","date":"2016-10-20","proceeding":null,"authors":["David Lopez-Paz","Maxime Oquab"],"abstract":"The goal of two-sample tests is to assess whether two samples, $S_P \\sim P^n$\nand $S_Q \\sim Q^m$, are drawn from the same distribution. Perhaps intriguingly,\none relatively unexplored method to build two-sample tests is the use of binary\nclassifiers. In particular, construct a dataset by pairing the $n$ examples in\n$S_P$ with a positive label, and by pairing the $m$ examples in $S_Q$ with a\nnegative label. If the null hypothesis \"$P = Q$\" is true, then the\nclassification accuracy of a binary classifier on a held-out subset of this\ndataset should remain near chance-level. As we will show, such Classifier\nTwo-Sample Tests (C2ST) learn a suitable representation of the data on the fly,\nreturn test statistics in interpretable units, have a simple null distribution,\nand their predictive uncertainty allow to interpret where $P$ and $Q$ differ.\nThe goal of this paper is to establish the properties, performance, and uses of\nC2ST. First, we analyze their main theoretical properties. Second, we compare\ntheir performance against a variety of state-of-the-art alternatives. Third, we\npropose their use to evaluate the sample quality of generative models with\nintractable likelihoods, such as Generative Adversarial Networks (GANs).\nFourth, we showcase the novel application of GANs together with C2ST for causal\ndiscovery.","url_abs":"http://arxiv.org/abs/1610.06545v4","url_pdf":"http://arxiv.org/pdf/1610.06545v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"revisiting-classifier-two-sample-tests","repo_url":"https://github.com/lopezpaz/classifier_tests","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"causal-discovery","task_name":"Causal Discovery"},{"task_slug":"two","task_name":"Vocal Bursts Valence Prediction"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1610.06545","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}