Browse State-of-the-Art › Two-sample testing
Two-sample testing
84 papers with code · 5 benchmarks · 1 dataset archive 2025-07-28
In statistical hypothesis testing, a two-sample test is a test performed on the data of two random samples, each independently obtained from a different given population. The purpose of the test is to determine whether the difference between these two populations is statistically significant. The statistics used in two-sample tests can be used to solve many machine learning problems, such as domain adaptation, covariate shift and generative adversarial networks.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
5 leaderboard tables shown for this task, 5 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| Blob (9 modes, 40 for each) (1 row) | MMD-D | Learning Deep Kernels for Non-Parametric Two-Sample Tests | code | Syntology ran 2 of 2 samples · 0 unverified | Compare |
| CIFAR-10 vs CIFAR-10.1 (1000 samples) (1 row) | MMD-D | Learning Deep Kernels for Non-Parametric Two-Sample Tests | code | Syntology ran 2 of 2 samples · 0 unverified | Compare |
| HDGM (d=10, N=4000) (1 row) | MMD-D | Learning Deep Kernels for Non-Parametric Two-Sample Tests | code | Syntology ran 2 of 2 samples · 0 unverified | Compare |
| HIGGS Data Set (1 row) | MMD-D | Learning Deep Kernels for Non-Parametric Two-Sample Tests | code | Syntology ran 2 of 2 samples · 0 unverified | Compare |
| MNIST vs Fake MNIST (1 row) | MMD-D | Learning Deep Kernels for Non-Parametric Two-Sample Tests | code | Syntology ran 2 of 2 samples · 0 unverified | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
1 dataset whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 84 papers with code (338 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
12 Dec 2017 7 repositories listedGenerative adversarial networks (GANs) are innovative techniques for learning generative models of complex data distributions from samples.
-
14 Dec 2018 5 repositories listedWe thus first propose a measure of `sensitivity' and show empirically that normal samples and adversarial samples have distinguishable sensitivity.
-
3 Jul 2019 4 repositories listedWe introduce hyppo, a unified library for performing multivariate hypothesis testing, including independence, two-sample, and k-sample testing.
-
9 Jun 2019 4 repositories listedBased on automatic deep learning segmentations, we extracted three features which quantify two-dimensional and three-dimensional characteristics of the tumors.
-
17 Jun 2022 3 repositories listed Syntology ran 10 of 11 samples · 1 unverified · 2 pointer-only (licence)Two-sample tests are important in statistics and machine learning, both as tools for scientific discovery as well as to detect distribution shifts.
-
28 Oct 2021 3 repositories listed Syntology ran 4 of 15 samples · 11 unverifiedIn practice, this parameter is unknown and, hence, the optimal MMD test with this particular kernel cannot be used.
-
7 May 2019 3 repositories listedMore precisely, the privacy guarantees of \emph{any} hypothesis testing based definition of privacy (including original DP) converges to GDP in the limit under composition.
-
10 Feb 2015 3 repositories listedWe consider the problem of learning deep generative models from data.
-
5 Oct 2020 2 repositories listed Syntology ran 4 of 8 samples · 4 unverified · 8 pointer-only (licence)To overcome this difficulty, we introduce a conditional selective inference (SI) framework -- a new statistical inference framework for data-driven hypotheses that has recently received considerable attention -- to…
-
23 Jun 2020 2 repositories listed Syntology ran 8 of 11 samples · 3 unverified · 11 pointer-only (licence)However, the adoption of such policies in practice is often challenging, as they are hard to interpret within the application context, and lack measures of uncertainty for the learned policy value and its decisions.
-
18 Jun 2020 2 repositories listed Syntology ran 5 of 20 samples · 15 unverifiedWe present an improved method for symbolic regression that seeks to fit data to formulas that are Pareto-optimal, in the sense of having the best accuracy for a given complexity.
-
24 Feb 2020 2 repositories listedIn this paper, we present ACORE (Approximate Computation via Odds Ratio Estimation), a frequentist approach to LFI that first formulates the classical likelihood ratio test (LRT) as a parametrized classification…
-
17 Feb 2020 2 repositories listedTo make decisions based on a model fit with auto-encoding variational Bayes (AEVB), practitioners often let the variational distribution serve as a surrogate for the posterior distribution.
-
19 Sep 2019 2 repositories listed Syntology ran 3 of 3 samples · 0 unverified · 3 pointer-only (licence)Here, we show that Lᵖ distances (with p≥1) between these distribution representatives give metrics on the space of distributions that are well-behaved to detect differences between distributions as they metrize the weak…
-
16 Apr 2019 2 repositories listedThroughout the last decade, random forests have established themselves as among the most accurate and popular supervised learning methods.
-
29 Oct 2018 2 repositories listed Syntology ran 2 of 6 samples · 4 unverifiedWe might hope that when faced with unexpected inputs, well-designed software systems would fire off warnings.
-
1 Aug 2018 2 repositories listedAs a result, a simple chi-squared test is obtained, where a test statistic depends on a mean and covariance of empirical differences between the samples, which we perturb for a privacy guarantee.
-
27 Feb 2017 2 repositories listedUnder Markovian assumptions, we leverage a Central Limit Theorem (CLT) for the empirical measure in the test statistic of the composite hypothesis Hoeffding test so as to establish weak convergence results for the test…
-
19 Feb 2017 2 repositories listedRobust PCA methods are typically batch algorithms which requires loading all observations into memory before processing.
-
13 Nov 2024 1 repository listedWe explore the trade-off between privacy and statistical utility in private two-sample testing under local differential privacy (LDP) for both multinomial and continuous data.
-
26 Oct 2024 1 repository listed Syntology ran 0 of 5 samples · 5 unverified · 5 pointer-only (licence)Users often interact with large language models through black-box inference APIs, both for closed- and open-weight models (e.
-
22 Oct 2024 1 repository listedOur first framework allows us to convert any conditional independence test into a conditional two-sample test in a black-box manner, while preserving the asymptotic properties of the original conditional independence…
-
16 Oct 2024 1 repository listed Syntology ran 12 of 12 samples · 0 unverifiedWe introduce credal two-sample testing, a new hypothesis testing framework for comparing credal sets -- convex sets of probability measures where each element captures aleatoric uncertainty and the set itself represents…
-
12 Sep 2024 1 repository listedThe focus of this study is to evaluate the effectiveness of Machine Learning (ML) methods for two-sample testing with right-censored observations.
-
12 Jul 2024 1 repository listedDespite its success and widespread adoption, the primary limitation of the MMD test has been its quadratic-time complexity, which poses challenges for large-scale analysis.
-
30 Oct 2023 1 repository listedWe propose a general framework for constructing powerful, sequential hypothesis tests for a large class of nonparametric testing problems.
-
17 Aug 2023 1 repository listedGiven n observations from two balanced classes, consider the task of labeling an additional m inputs that are known to all belong to \emph{one} of the two classes.
-
14 Jun 2023 1 repository listed Syntology ran 2 of 15 samples · 13 unverifiedWe propose novel statistics which maximise the power of a two-sample test based on the Maximum Mean Discrepancy (MMD), by adapting over the set of kernels used in defining it.
-
25 May 2023 1 repository listedWe then compute the Jensen-Shannon divergence between these operators, thereby establishing a proper divergence measure between probability distributions in the input space.
-
8 Mar 2023 1 repository listedMachine learning and deep learning have been used extensively to classify physical surfaces through images and time-series contact data.
Syntology lines on 10 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections