Browse State-of-the-Art › Dataset Distillation
Dataset Distillation
120 papers with code · 0 benchmarks · 0 datasets archive 2025-07-28
Dataset distillation is the task of synthesizing a small dataset such that models trained on it achieve high performance on the original large dataset. A dataset distillation algorithm takes as input a large real dataset to be distilled (training set), and outputs a small synthetic distilled dataset, which is evaluated via testing models trained on this distilled dataset on a separate real dataset (validation/test set). A good small distilled dataset is not only useful in dataset understanding, but has various applications (e.g., continual learning, privacy, neural architecture search, etc.).
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
No dataset record in the archive lists this task.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 120 papers with code (216 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
22 Mar 2022 6 repositories listedTo efficiently obtain the initial and target network parameters for large-scale datasets, we pre-compute and store training trajectories of expert networks trained on the real dataset.
-
27 Nov 2018 6 repositories listed Syntology ran 3 of 17 samples · 14 unverified · 3 pointer-only (licence)Model distillation aims to distill the knowledge of a complex model into a simpler one.
-
6 Dec 2023 4 repositories listed Syntology ran 3 of 10 samples · 7 unverified · 4 pointer-only (licence)Contemporary machine learning requires training large neural networks on massive datasets and thus faces the challenges of high computational demands.
-
22 May 2024 3 repositories listedFedCache 2.
-
20 Nov 2022 3 repositories listedTo mitigate the adverse impact of this accumulated trajectory error, we propose a novel approach that encourages the optimization algorithm to seek a flat trajectory.
-
30 Oct 2022 3 repositories listed Syntology ran 0 of 7 samples · 7 unverifiedIn this paper, we study \xw{dataset distillation (DD)}, from a novel perspective and introduce a \emph{dataset factorization} approach, termed \emph{HaBa}, which is a plug-and-play strategy portable to any existing DD…
-
27 Jul 2021 3 repositories listedThe effectiveness of machine learning algorithms arises from being able to extract useful features from large amounts of data.
-
6 Oct 2019 3 repositories listed Syntology ran 0 of 13 samples · 13 unverifiedWe propose to simultaneously distill both images and their labels, thus assigning each synthetic sample a `soft' label (a distribution of labels).
-
30 Jun 2025 2 repositories listed Syntology ran 1 of 2 samples · 1 unverified · 2 pointer-only (licence)In this study, we integrate vision-language methods into DD by introducing text prototypes to distill language information and collaboratively synthesize data with image prototypes, thereby enhancing dataset…
-
16 Aug 2024 2 repositories listed Syntology ran 2 of 2 samples · 0 unverifiedIn this study, we proposed a novel generative dataset distillation method based on Stable Diffusion.
-
13 Nov 2023 2 repositories listed Syntology ran 1 of 1 samples · 0 unverifiedRe-examining the foundational back-propagation through time method, we study the pronounced variance in the gradients, computational burden, and long-term dependencies.
-
13 Oct 2023 2 repositories listed Syntology ran 8 of 9 samples · 1 unverifiedWe validate the proposed SGDD across 9 datasets and achieve state-of-the-art results on all of them: for example, on the YelpChi dataset, our approach maintains 98.
-
10 Oct 2023 2 repositories listed Syntology ran 2 of 3 samples · 1 unverifiedTo achieve this, we also introduce the MSE between representations of the inner model and the self-supervised target model on the original full dataset for outer optimization.
-
29 Sep 2023 2 repositories listed Syntology ran 35 of 43 samples · 8 unverified · 43 pointer-only (licence)Emerging research on dataset distillation aims to reduce training costs by creating a small synthetic set that contains the information of a larger real dataset and ultimately achieves test accuracy equivalent to a…
-
15 Aug 2023 2 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)In this work, we design the first vision-language dataset distillation method, building on the idea of trajectory matching.
-
18 Jul 2023 2 repositories listedBy distilling both InD samples and outliers, the condensed datasets are capable of training models competent in both InD classification and OOD detection.
-
22 Jun 2023 2 repositories listed Syntology ran 1 of 6 samples · 5 unverified · 6 pointer-only (licence)The proposed method demonstrates flexibility across diverse dataset scales and exhibits multiple advantages in terms of arbitrary resolutions of synthesized images, low training cost and memory consumption with…
-
28 May 2023 2 repositories listed Syntology ran 1 of 8 samples · 7 unverifiedWe believe this paradigm will open up new avenues in the dynamics of distillation and pave the way for efficient dataset distillation.
-
2 May 2023 2 repositories listedDataset Distillation aims to distill an entire dataset's knowledge into a few synthetic images.
-
8 Mar 2023 2 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)To the best of our knowledge, we are the first to achieve higher accuracy on complex architectures than simple ones, such as 75.
-
28 Feb 2023 2 repositories listedAlthough there are various matching objectives, currently the strategy for selecting original images is limited to naive random sampling.
-
13 Feb 2023 2 repositories listedWe propose a new dataset distillation algorithm using reparameterization and convexification of implicit gradients (RCIG), that substantially improves the state-of-the-art.
-
3 Jan 2023 2 repositories listed Syntology ran 6 of 15 samples · 9 unverifiedA model trained on this smaller distilled dataset can attain comparable performance to a model trained on the original training dataset.
-
12 Dec 2022 2 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)Dataset Distillation (DD), a newly emerging field, aims at generating much smaller but efficient synthetic training datasets from large ones.
-
19 Nov 2022 2 repositories listed Syntology ran 0 of 2 samples · 2 unverifiedThe resulting algorithm sets new SOTA on ImageNet-1K: we can scale up to 50 IPCs (Image Per Class) on ImageNet-1K on a single GPU (all previous methods can only scale to 2 IPCs on ImageNet-1K), leading to the best…
-
21 Oct 2022 2 repositories listed Syntology ran 1 of 2 samples · 1 unverified · 2 pointer-only (licence)Dataset distillation compresses large datasets into smaller synthetic coresets which retain performance with the aim of reducing the storage and computational burden of processing the entire dataset.
-
24 Aug 2022 2 repositories listed Syntology ran 1 of 5 samples · 4 unverifiedIn federated learning, all networked clients contribute to the model training cooperatively.
-
6 Jun 2022 2 repositories listedWe propose an algorithm that compresses the critical information of a large dataset into compact addressable memories.
-
1 Jun 2022 2 repositories listed Syntology ran 0 of 1 samples · 1 unverified · 1 pointer-only (licence)Dataset distillation can be formulated as a bi-level meta-learning problem where the outer loop optimizes the meta-dataset and the inner loop trains a model on the distilled data.
-
15 Jun 2020 2 repositories listed Syntology ran 12 of 14 samples · 2 unverifiedIn particular, we study the problem of label distillation - creating synthetic labels for a small set of real images, and show it to be more effective than the prior image-based approach to dataset distillation.
Syntology lines on 21 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections