Browse State-of-the-Art › Dataset Condensation
Dataset Condensation
35 papers with code · 0 benchmarks · 0 datasets archive 2025-07-28
Condense the full dataset into a tiny set of synthetic data.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
No dataset record in the archive lists this task.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 35 papers with code (56 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
10 Jun 2020 5 repositories listedAs the state-of-the-art machine learning methods in many fields rely on larger datasets, storing datasets and training models on them become significantly more expensive.
-
8 Oct 2021 4 repositories listedComputational cost of training state-of-the-art deep models in many learning problems is rapidly increasing due to more sophisticated models and larger datasets.
-
29 Nov 2023 3 repositories listed Syntology ran 5 of 7 samples · 2 unverified · 7 pointer-only (licence)We call this perspective "generalized matching" and propose Generalized Various Backbone and Statistical Matching (G-VBSM) in this work, which aims to create a synthetic dataset with densities, ensuring consistency with…
-
15 Jun 2022 3 repositories listed Syntology ran 1 of 4 samples · 3 unverifiedHowever, existing approaches have their inherent limitations: (1) they are not directly applicable to graphs where the data is discrete; and (2) the condensation process is computationally expensive due to the involved…
-
15 Apr 2022 3 repositories listed Syntology ran 3 of 7 samples · 4 unverified · 2 pointer-only (licence)However, traditional GANs generated images are not as informative as the real training samples when being used to train deep neural networks.
-
26 Dec 2023 2 repositories listed Syntology ran 1 of 1 samples · 0 unverifiedNowadays, optimization-oriented methods have been the primary method in the field of dataset condensation for achieving SOTA results.
-
19 Jul 2023 2 repositories listed Syntology ran 6 of 13 samples · 7 unverified · 13 pointer-only (licence)In this paper, we propose a novel dataset condensation method based on distribution matching, which is more efficient and promising.
-
22 Jun 2023 2 repositories listed Syntology ran 1 of 6 samples · 5 unverified · 6 pointer-only (licence)The proposed method demonstrates flexibility across diverse dataset scales and exhibits multiple advantages in terms of arbitrary resolutions of synthesized images, low training cost and memory consumption with…
-
21 Oct 2022 2 repositories listed Syntology ran 1 of 2 samples · 1 unverified · 2 pointer-only (licence)Dataset distillation compresses large datasets into smaller synthetic coresets which retain performance with the aim of reducing the storage and computational burden of processing the entire dataset.
-
20 Jul 2022 2 repositories listed Syntology ran 4 of 7 samples · 3 unverifiedDataset Condensation is a newly emerging technique aiming at learning a tiny dataset that captures the rich information encoded in the original dataset.
-
30 May 2022 2 repositories listed Syntology ran 1 of 1 samples · 0 unverifiedThe great success of machine learning with massive amounts of data comes at a price of huge computation costs and storage for training and tuning.
-
3 Mar 2022 2 repositories listed Syntology ran 27 of 33 samples · 6 unverified · 33 pointer-only (licence)Dataset condensation aims at reducing the network training effort through condensing a cumbersome training set into a compact synthetic one.
-
7 Feb 2022 2 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)However, in this study, we prove that the existing DC methods can perform worse than the random selection method when task-irrelevant information forms a significant part of the training dataset.
-
16 Feb 2021 2 repositories listedIn many machine learning problems, large-scale datasets have become the de-facto standard to train state-of-the-art deep networks at the price of heavy computation load.
-
30 Dec 2024 1 repository listed Syntology ran 3 of 17 samples · 14 unverifiedSpecifically, our work delves into three key aspects to provide valuable empirical insights: (1) temporal processing of video data, (2) establishing a comprehensive evaluation protocol for video dataset condensation,…
-
23 Dec 2024 1 repository listedThe main bottleneck of the above paradigms is whether the effective information of the original graph is fully preserved when consenting to the primary sub-scale (the first of multiple scales), which determines the…
-
6 Dec 2024 1 repository listedConventionally, DC relies on a costly bi-level optimization which prohibits its practicality.
-
22 Sep 2024 1 repository listed Syntology ran 6 of 9 samples · 3 unverified · 9 pointer-only (licence)In response to this limitation, we introduce a novel method, Heterogeneous Model Dataset Condensation (HMDC), designed to produce universally applicable condensed images through cross-model interactions.
-
6 Jun 2024 1 repository listedDataset Distillation has emerged as a technique for compressing large datasets into smaller synthetic counterparts, facilitating downstream training tasks.
-
4 Jun 2024 1 repository listed Syntology ran 0 of 1 samples · 1 unverified · 1 pointer-only (licence)To mitigate this gap, we theoretically analyze the optimization objective of dataset condensation for TS-forecasting and propose a new one-line plugin of dataset condensation designated as Dataset Condensation for Time…
-
3 Jun 2024 1 repository listed Syntology ran 14 of 18 samples · 4 unverified · 18 pointer-only (licence)Specifically, from the inner-class view, we construct multiple "middle encoders" to perform pseudo long-term distribution alignment, making the condensed set a good proxy of the real one during the whole training…
-
21 Apr 2024 1 repository listed Syntology ran 1 of 2 samples · 1 unverified · 1 pointer-only (licence)Dataset condensation, a concept within data-centric learning, efficiently transfers critical attributes from an original dataset to a synthetic version, maintaining both diversity and realism.
-
18 Mar 2024 1 repository listedCurrent methods frame this as maximizing the distilled classification accuracy for a budget of K distilled images-per-class, where K is a positive integer.
-
12 Mar 2024 1 repository listed Syntology ran 5 of 9 samples · 4 unverified · 9 pointer-only (licence)Different from previous methods, our proposed framework aims to generate a condensed dataset that matches the surrogate objectives in both the time and frequency domains.
-
10 Mar 2024 1 repository listed Syntology ran 15 of 19 samples · 4 unverifiedThese two challenges connect to the "subset degradation problem" in traditional dataset condensation: a subset from a larger condensed dataset is often unrepresentative compared to directly condensing the whole dataset…
-
8 Feb 2024 1 repository listedThis synthetic dataset retains the essential information of the original dataset, enabling models trained on it to achieve performance levels comparable to those trained on the full dataset.
-
31 Jan 2024 1 repository listedTo achieve this goal, we propose new dataset condensation techniques and an innovative unlearning scheme that strikes a balance between machine unlearning privacy, utility, and efficiency.
-
21 Oct 2023 1 repository listed Syntology ran 12 of 16 samples · 4 unverified · 16 pointer-only (licence)However, these scenarios have two significant challenges: 1) the varying computational resources available on the devices require a dataset size different from the pre-defined condensed dataset, and 2) the limited…
-
17 Oct 2023 1 repository listed Syntology ran 6 of 7 samples · 1 unverified · 7 pointer-only (licence)The rapid development of Internet technology has given rise to a vast amount of graph-structured data.
-
2 Oct 2023 1 repository listedSpecifically, we model the discrete user-item interactions via a probabilistic approach and design a pre-augmentation module to incorporate the potential preferences of users into the condensed datasets.
Syntology lines on 20 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections