Browse State-of-the-Art › Continual Pretraining
Continual Pretraining
34 papers with code · 3 benchmarks · 3 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
3 leaderboard tables shown for this task, 3 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| ACL-ARC (1 row) | DAS | Continual Pre-training of Language Models | code | — | Compare |
| AG News (1 row) | CPT | Continual Training of Language Models for Few-Shot Learning | code | — | Compare |
| SciERC (1 row) | DAS | Continual Pre-training of Language Models | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
3 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 34 papers with code (70 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
11 Apr 2024 3 repositories listedUnlike traditional LMs that learn to predict every next token in a corpus, Rho-1 employs Selective Language Modeling (SLM), which selectively trains on useful tokens that aligned with the desired distribution.
-
11 Oct 2022 3 repositories listedRecent work on applying large language models (LMs) achieves impressive performance in many NLP applications.
-
12 Feb 2024 2 repositories listed Syntology ran 3 of 3 samples · 0 unverified · 3 pointer-only (licence)We release our curated AutoMathText dataset to facilitate future research in automated domain-specific data curation.
-
27 Sep 2023 2 repositories listedWe also examine the impact of various design choices in the pretraining process, including the data mix and the training curriculum of sequence lengths -- our ablation experiments suggest that having abundant long texts…
-
9 Feb 2023 2 repositories listed Syntology ran 2 of 6 samples · 4 unverifiedGeospatial technologies are becoming increasingly essential in our world for a wide range of applications, including agriculture, urban planning, and disaster response.
-
7 Feb 2023 2 repositories listedA novel proxy is also proposed to preserve the general knowledge in the original LM.
-
30 Dec 2020 2 repositories listedWhile pre-trained language models (PTLMs) have achieved noticeable success on many NLP tasks, they still struggle for tasks that require event temporal reasoning, which is essential for event-centric applications.
-
31 Jul 2020 2 repositories listedIn this paper, we challenge this assumption by showing that for domains with abundant unlabeled text, such as biomedicine, pretraining language models from scratch results in substantial gains over continual pretraining…
-
30 May 2025 1 repository listedThe performance of large language models (LLMs) continues to improve, as reflected in rising scores on standard benchmarks.
-
22 May 2025 1 repository listedWe present a Japanese domain-specific language model for the pharmaceutical field, developed through continual pretraining on 2 billion Japanese pharmaceutical tokens and 8 billion English biomedical tokens.
-
2 Apr 2025 1 repository listedLarge Language Models (LLMs) trained on historical web data inevitably become outdated.
-
6 Mar 2025 1 repository listedOur watermarks are designed to be memorized by the LLM through seamlessly integrating in its training data, making them harder to detect lexically during preprocessing.
-
9 Jan 2025 1 repository listed Syntology ran 4 of 13 samples · 9 unverified · 13 pointer-only (licence)Domain-adaptive post-training of large language models (LLMs) has emerged as a promising approach for specialized domains such as medicine and finance.
-
11 Dec 2024 1 repository listedThe integration of artificial intelligence (AI) in legal judgment prediction (LJP) has the potential to transform the legal landscape, particularly in jurisdictions like India, where a significant backlog of cases…
-
21 Oct 2024 1 repository listedExperimental results demonstrate the effectiveness of our approach, achieving a 5% absolute performance improvement on Leandojo benchmark.
-
26 Sep 2024 1 repository listedHowever, this removal increases the burden on token embeddings to encode all language-specific information, which may hinder the model's ability to produce more language-neutral representations.
-
26 Aug 2024 1 repository listedIn this work, we complement current perspectives on continual pretraining through a research test bed as well as provide comprehensive guidance for effective continual model updates in such scenarios.
-
18 Jul 2024 1 repository listedThis paper introduces long-context Granite code models that support effective context windows of up to 128K tokens.
-
10 Jun 2024 1 repository listedThis survey delves into the sophisticated landscape of lifelong learning, categorizing strategies into two primary groups: Internal Knowledge and External Knowledge.
-
30 May 2024 1 repository listedSecond, we revisit and explore cross-domain continual pretraining for both multispectral and SAR imagery, building efficient EO foundation models from strongest vision models such as DINOv2.
-
20 May 2024 1 repository listed Syntology ran 1 of 6 samples · 5 unverifiedLow-rank adaptation is a popular parameter-efficient fine-tuning method for large language models.
-
24 Apr 2024 1 repository listed Syntology ran 7 of 7 samples · 0 unverified · 7 pointer-only (licence)Despite the recent progress in long-context language models, it remains elusive how transformer-based models exhibit the capability to retrieve relevant information from arbitrary locations within the long context.
-
7 Mar 2024 1 repository listed Syntology ran 2 of 8 samples · 6 unverifiedThe Yi model family is based on 6B and 34B pretrained language models, then we extend them to chat models, 200K long context models, depth-upscaled models, and vision-language models.
-
15 Feb 2024 1 repository listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)We demonstrate that continual pretraining of the full model on 1B-5B tokens of such data is an effective and affordable strategy for scaling the context length of language models to 128K.
-
11 Nov 2023 1 repository listed Syntology ran 3 of 3 samples · 0 unverified · 3 pointer-only (licence)The limited availability of labelled data in Action Quality Assessment (AQA), has forced previous works to fine-tune their models pretrained on large-scale domain-general datasets.
-
14 Aug 2023 1 repository listed Syntology ran 10 of 14 samples · 4 unverifiedRegarding the growing nature of real-world data, such an offline training paradigm on ever-expanding data is unsustainable, because models lack the continual learning ability to accumulate knowledge constantly.
-
1 Jan 2023 1 repository listedRegarding the growing nature of real-world data, such an offline training paradigm on ever-expanding data is unsustainable, because models lack the continual learning ability to accumulate knowledge constantly.
-
21 Nov 2022 1 repository listedContinual pretraining is a popular way of building a domain-specific pretrained language model from a general-domain language model.
-
8 Nov 2022 1 repository listedWe conducted experiments using our method on datasets with a large vocabulary gap from a source domain.
-
19 May 2022 1 repository listed Syntology ran 0 of 7 samples · 7 unverifiedWe formalize and investigate the characteristics of the continual pre-training scenario in both language and vision environments, where a model is continually pre-trained on a stream of incoming data and only later…
Syntology lines on 10 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections