Browse State-of-the-Art › Memorization
Memorization
438 papers with code · 1 benchmark · 4 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
1 leaderboard table shown for this task, 1 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| BIG-bench (Hindu Knowledge) (3 rows) | PaLM-540B (few-shot, k=5) | PaLM: Scaling Language Modeling with Pathways | code | Syntology ran 30 of 37 samples · 7 unverified | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
4 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 438 papers with code (1,088 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
25 Oct 2017 71 repositories listed Syntology ran 30 of 47 samples · 17 unverified · 15 pointer-only (licence)We also find that mixup reduces the memorization of corrupt labels, increases the robustness to adversarial examples, and stabilizes the training of generative adversarial networks.
-
24 Jun 2016 39 repositories listed Syntology ran 0 of 5 samples · 5 unverified · 5 pointer-only (licence)Memorization of feature interactions through a wide set of cross-product feature transformations are effective and interpretable, while generalization requires more feature engineering effort.
-
31 Oct 2016 11 repositories listedThe ByteNet is a one-dimensional convolutional neural network that is composed of two parts, one to encode the source sequence and the other to decode the target sequence.
-
5 Apr 2022 7 repositories listed Syntology ran 30 of 37 samples · 7 unverifiedTo further our understanding of the impact of scale on few-shot learning, we trained a 540-billion parameter, densely activated, Transformer language model, which we call Pathways Language Model PaLM.
-
9 Jun 2022 6 repositories listed Syntology ran 0 of 3 samples · 3 unverifiedBIG-bench focuses on tasks that are believed to be beyond the capabilities of current language models.
-
6 Jan 2022 6 repositories listed Syntology ran 2 of 14 samples · 12 unverifiedIn this paper we propose to study generalization of neural networks on small algorithmically generated datasets.
-
1 Nov 2019 5 repositories listed Syntology ran 3 of 3 samples · 0 unverifiedApplying this augmentation to a strong Wikitext-103 LM, with neighbors drawn from the original training set, our $k$NN-LM achieves a new state-of-the-art perplexity of 15.
-
18 Apr 2018 5 repositories listed Syntology ran 7 of 7 samples · 0 unverified · 7 pointer-only (licence)Deep learning with noisy labels is practically challenging, as the capacity of deep models is so high that they can totally memorize these noisy labels sooner or later during training.
-
Exposing flaws of generative model evaluation metrics and their unfair treatment of diffusion models7 Jun 2023 4 repositories listed Syntology ran 30 of 63 samples · 33 unverified · 1 pointer-only (licence)Comparing to 17 modern metrics for evaluating the overall performance, fidelity, diversity, rarity, and memorization of generative models, we find that the state-of-the-art perceptual realism of diffusion models as…
-
3 Apr 2023 4 repositories listedHow do large language models (LLMs) develop and evolve over the course of training?
-
22 Oct 2021 4 repositories listedThese observations require us to rethink the treatment of noisy labels, and we hope the availability of these two datasets would facilitate the development and evaluation of future learning with noisy label solutions.
-
1 Sep 2020 4 repositories listed Syntology ran 0 of 5 samples · 5 unverifiedIn this paper, we investigate the principle that `good explanations are hard to vary' in the context of deep learning.
-
14 May 2019 4 repositories listed Syntology ran 5 of 22 samples · 17 unverified · 9 pointer-only (licence)In this article, we first frame the research problem of optimizing an adaptive and personalized spaced repetition scheduler when memorization concerns the application of underlying multiple skills.
-
22 May 2025 3 repositories listed Syntology ran 1 of 10 samples · 9 unverifiedIn this paper, we introduce R1-Searcher++, a novel framework designed to train LLMs to adaptively leverage both internal and external knowledge sources.
-
5 Feb 2025 3 repositories listedWhile conventional wisdom suggests that sophisticated reasoning tasks demand extensive training data (>100, 000 examples), we demonstrate that complex mathematical reasoning abilities can be effectively elicited with…
-
8 Jul 2024 3 repositories listedData owners may request the removal of their data from a trained model due to privacy or copyright concerns.
-
17 Jan 2022 3 repositories listedMachine Learning models face increased concerns regarding the storage of personal user data and adverse impacts of corrupted data like backdoors or systematic bias.
-
8 Dec 2021 3 repositories listedLanguage modelling provides a step towards intelligent communication systems by harnessing large repositories of written human knowledge to better predict and understand the world.
-
10 Jul 2021 3 repositories listedA dynamic transition mechanism is used to move from supervision loss in early learning to consistency loss for consensus of predictions among networks in the later stage.
-
6 Nov 2019 3 repositories listed Syntology ran 0 of 1 samples · 1 unverified · 1 pointer-only (licence)Sample selection approaches are popular in robust learning from noisy labels.
-
14 Jan 2019 3 repositories listedLearning with noisy labels is one of the hottest problems in weakly-supervised learning.
-
9 Feb 2016 3 repositories listedWe investigate a new method to augment recurrent neural networks with extra memory without increasing the number of network parameters.
-
14 Apr 2025 2 repositories listed Syntology ran 4 of 4 samples · 0 unverifiedScientific equation discovery is a fundamental task in the history of scientific progress, enabling the derivation of laws governing natural phenomena.
-
31 Dec 2024 2 repositories listedThis paper introduces a Hierarchical visual token Compression (HiCo) method designed for high-fidelity representation and a practical context modeling system VideoChat-Flash tailored for multimodal long-sequence…
-
30 Oct 2024 2 repositories listed Syntology ran 3 of 14 samples · 11 unverifiedWarm-starting neural network training by initializing networks with previously learned weights is appealing, as practical neural networks are often deployed under a continuous influx of new data.
-
17 Jun 2024 2 repositories listed Syntology ran 7 of 7 samples · 0 unverified · 3 pointer-only (licence)First, counterintuitively, we observe that pretraining on more data shows no significant improvement in the model's capability to acquire and maintain factual knowledge.
-
27 Dec 2023 2 repositories listed Syntology ran 5 of 6 samples · 1 unverifiedIn a simple setting with direct supervision on the generative factors, we show how learning class-agnostic transformations offers a way to circumvent catastrophic forgetting and improve classification accuracy over time.
-
10 Dec 2023 2 repositories listedSince clean samples are easier distinguished by GMM with increasing noise, the memory bank can still maintain high quality at a high noise ratio.
-
10 Nov 2023 2 repositories listed Syntology ran 3 of 3 samples · 0 unverifiedWe propose the Data Contamination Quiz (DCQ), a simple and effective approach to detect data contamination in large language models (LLMs) and estimate the amount of it.
-
10 Nov 2023 2 repositories listed Syntology ran 5 of 6 samples · 1 unverified · 6 pointer-only (licence)However, this hypothesis heavily relies on the overfitting of target models, which will be mitigated by multiple regularization methods and the generalization of LLMs.
Syntology lines on 18 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections