Papers › Rethinking Dataset Distillation for Classification: Do Distilled Sets Outperform Coresets?
Rethinking Dataset Distillation for Classification: Do Distilled Sets Outperform Coresets?
Trisha Mittal, Akshay Mehra, Joshua Kimball
Title, abstract, authors and date from arXiv's metadata (CC0); this paper is not in the Papers with Code archive (frozen 2025-07-28).
Dataset distillation (DD) has emerged as a prominent approach in data centric machine learning, aiming to synthesize compact training sets for efficient training by compressing the information in large datasets into a small number of synthetic samples. However, DD methods are often evaluated under inconsistent evaluation protocols, ranging from standard ERM to single/multi-teacher supervision, making it difficult to isolate the effectiveness of distilled data from evaluation. Moreover, many prior methods claim that DD outperforms data pruning approaches such as coreset selection (CS), based on the assumption that restricting condensed datasets to subsets of real samples fundamentally limits their expressiveness. In this work, we critically evaluate DD methods through large-scale experiments using standardized datasets and evaluation protocols to assess their intrinsic effectiveness. We benchmark seven state-of-the-art (SOTA) DD methods on ImageNet-1K, ImageNet100, and ImageNette, using three widely adopted training protocols against three CS strategies. Our results show that while some DD methods fail to outperform even simple random subsets, the SOTA DD approaches are comparable to or worse than coresets on large-scale datasets and incur a substantially higher cost for construction. Beyond accuracy, we also evaluate the representativeness, diversity, and quality of condensed sets, and find that coresets consistently achieve better coverage of the original data distribution. These findings highlight the limited practical advantages of current DD methods and show that coresets remain competitive and are often a more computationally efficient alternative for data-centric learning.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
For agents, Syntology's MCP tool lists every function and class Syntology harvested from this paper and whether it ran (how to connect): get_harvested_code_for_paper(arxiv_id="2606.18209")
Code
Syntology Ran 61 of 72 code samples harvested from 5 repositories linked to this paper; 11 have no recorded run. Of those that ran: 5 ran · honoured contract; 1 ran · violated contract; 15 ran · our draft was wrong; 40 ran with no contract checked.
By repository: found in paper text by Syntology: 72 samples from 5 repositories, 61 ran. The run record, sample by sample. “Ran” means executed on a synthesized input, not that the code is correct or reproduces the paper.
Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
72 samples harvested; 61 ran; 5 honoured the contract we drafted; 11 have no recorded run. Read from Syntology's graph 2026-09-24; that is when this build read the record, not when the samples ran.
Licence: 54 of the 72 samples are pointer only, meaning Syntology does not serve that copy's text. This page shows no code text for any sample; each one links to its file in the repository.
Harvested from 5 repositories linked to this paper, official or community; each sample names its own and says which. “Ran” means the sample executed on a synthesized input. It does not mean the output is correct, and nothing here reproduces the paper's results. “Honoured” and “violated” refer to a contract Syntology drafted from the code itself; “our draft was wrong” and “fixture could not drive it” are failures of Syntology's instrument, not of the code.
Each sample ends with its code_sha256, Syntology's identity for that exact code. An agent fetches the stored sample with Syntology's MCP tool get_code(code_sha256="…") (how to connect); click an identity to copy that call.
Repository labels, per sample. official repository: The archive marks this repository official for the paper. named in the paper: The archive records that the paper mentions this repository; it is not marked official. community (archive-listed): In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper. found in paper text by Syntology: Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted. community: Not in the archive's code links for this paper; a community repository Syntology harvested. Samples from a repository marked official are listed first. Licence labels name the repository's licence as recorded at harvest. “Pointer only” means Syntology does not serve that copy's text, for one of four reasons: no licence file was found; the licence was not identified; the licence is recorded as permissive but that copy's record is not marked cleared; or the licence is outside the permissive list Syntology serves text under (MIT, Apache-2.0, BSD and similar). Some licences outside that list permit redistribution, such as WTFPL, and GPL-3.0 under its conditions; they are simply not on the list. Hover a licence label for the reason. File links open the file on GitHub at the default branch, which may have changed since the harvest.
19442a6a71144dbd · report
3da7ee44118b27ad · report
96a2132550304d0d · report
339dfa096b3c4a02 · report
8937ae0759fe135c · report
e28ebe17022741b6 · report
c2749ad43673e986 · report
f9e75567dfd01367 · report
71310b3a95d5c6eb · report
77c5c0238ecc8b28 · report
7eb2592bfc829f81 · report
d6a68e210556f857 · report
fde7881344a30b12 · report
1712a07966b542ee · report
6fc038f89ff59f28 · report
dfaf3be9474cef03 · report
ab1c9568b4e13899 · report
fac5364e2f53c6db · report
a0131fb70c267a9e · report
408e747e0425a594 · report
1b4af1f0c7b88a83 · report
e157ae27e2ac5b4a · report
ca73dd5ac9ee210f · report
b07a34e357aed3db · report
dea99904c6a5b594 · report
734076d456219648 · report
b97000ad1f1f7bba · report
c92c27c924b517e8 · report
370445b84099fd37 · report
3e0fa4efc22272d4 · report
6b9edc653f0146c3 · report
a62a3ace38a10be2 · report
31acc8c84b0a2549 · report
ba1399ac9cd62aa7 · report
8e9046371d41825d · report
36e30c7fb679ec78 · report
25dda4be04d5c390 · report
d923e828eed4fb87 · report
cd4d8e43e2ff286b · report
6a59f07f0b669e9e · report
6728d2ca5cbb846d · report
585d0c5757aa75ee · report
9ce85b699efd5492 · report
38fd828e9ae4636a · report
580747f8e67364f5 · report
f6d7c009a8efb8b7 · report
bf057d506488cd4c · report
03310bba324ae4fb · report
8afbfc42c6ea0448 · report
8aac35516308515a · report
8ba16d20cf03f09b · report
0756a99508caf42c · report
bea2d776a332f2b0 · report
bb350b00da2aea93 · report
0b56680ee1708fee · report
ef1eb484f67917c1 · report
fc06c170864d1f3d · report
23bf22274cc64dc6 · report
4f312e21745a6f61 · report
52605edc9d326455 · report
fb78e314a9cbf5ff · report
57a1842709f61195 · report
6a0d5ecd441905cc · report
29947a0a94157558 · report
665d8a4e8f673a4c · report
fd6a20859bd2e457 · report
c9efd67365a73165 · report
d1d74d9fb90d59c1 · report
e21d40497e4a97b0 · report
5b2e2797a6da2cb5 · report
4a3eef12c129ad96 · report
230fb2e9963063c9 · report
Results from the paper
The Papers with Code archive ends with its 2025-07-28 snapshot. This paper's arXiv identifier, 2606.18209, was issued in June 2026, after that date, so the archive has no leaderboard rows for it.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections