Methods › Natural Language Processing › Language Models › Chinchilla
Chinchilla
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
Chinchilla is a 70B parameters model trained as a compute-optimal model with 1.4 trillion tokens. Findings suggest that these types of models are trained optimally by equally scaling both model size and training tokens. It uses the same compute budget as Gopher but with 4x more training data. Chinchilla and Gopher are trained for the same number of FLOPs. It is trained using MassiveText using a slightly modified SentencePiece tokenizer. More architectural details in the paper.
Papers archive 2025-07-28
25 shown of 25, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
Superposition Yields Robust Neural Scaling 15 May 2025 · 1 repository · arXiv:2505.10465
-
Compute-Optimal LLMs Provably Generalize Better With Scale 21 Apr 2025 · 0 repositories · arXiv:2504.15208
-
Scaling Inference-Efficient Language Models 30 Jan 2025 · 0 repositories · arXiv:2501.18107
-
Physics of Skill Learning 21 Jan 2025 · 1 repository · arXiv:2501.12391
-
The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws 21 Jan 2025 · 0 repositories · arXiv:2501.12486
-
How Much Can We Forget about Data Contamination? 4 Oct 2024 · 1 repository · arXiv:2410.03249Syntology ran 1 of 2 samples · 1 unverified
-
Energy Estimation of Last Mile Electric Vehicle Routes 21 Aug 2024 · 0 repositories · arXiv:2408.12006
-
Scaling Law with Learning Rate Annealing 20 Aug 2024 · 0 repositories · arXiv:2408.11029
-
Time Matters: Scaling Laws for Any Budget 27 Jun 2024 · 0 repositories · arXiv:2406.18922
-
Reconciling Kaplan and Chinchilla Scaling Laws 12 Jun 2024 · 1 repository · arXiv:2406.12907Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)
-
Revisiting MoE and Dense Speed-Accuracy Comparisons for LLM Training 23 May 2024 · 1 repository · arXiv:2405.15052
-
More Compute Is What You Need 30 Apr 2024 · 0 repositories · arXiv:2404.19484
-
MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies 9 Apr 2024 · 3 repositories · arXiv:2404.06395Syntology ran 1 of 1 samples · 0 unverified
-
VBART: The Turkish LLM 2 Mar 2024 · 0 repositories · arXiv:2403.01308
-
A Resource Model For Neural Scaling Law 7 Feb 2024 · 0 repositories · arXiv:2402.05164
-
Beyond Chinchilla-Optimal: Accounting for Inference in Language Model Scaling Laws 31 Dec 2023 · 0 repositories · arXiv:2401.00448
-
The Falcon Series of Open Language Models 28 Nov 2023 · 0 repositories · arXiv:2311.16867
-
Language Modeling Is Compression 19 Sep 2023 · 1 repository · arXiv:2309.10668Syntology ran 6 of 8 samples · 2 unverified
-
Does Circuit Analysis Interpretability Scale? Evidence from Multiple Choice Capabilities in Chinchilla 18 Jul 2023 · 0 repositories · arXiv:2307.09458
-
Tune As You Scale: Hyperparameter Optimization For Compute Efficient Training 13 Jun 2023 · 0 repositories · arXiv:2306.08055
-
Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster 6 Apr 2023 · 2 repositories · arXiv:2304.03208
-
Accelerating Large Language Model Decoding with Speculative Sampling 2 Feb 2023 · 5 repositories · arXiv:2302.01318
-
An Information-Theoretic Analysis of Compute-Optimal Neural Scaling Laws 2 Dec 2022 · 0 repositories · arXiv:2212.01365
-
Galactica: A Large Language Model for Science 16 Nov 2022 · 1 repository · arXiv:2211.09085Syntology ran 0 of 2 samples · 2 unverified
-
Training Compute-Optimal Large Language Models 29 Mar 2022 · 2 repositories · arXiv:2203.15556Syntology ran 8 of 11 samples · 3 unverified · 4 pointer-only (licence)
Tasks archive 2025-07-28
20 shown of 106 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections