Methods › Natural Language Processing › Language Models › Pythia
Pythia
Introduced by Stella Biderman et al. in Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
Pythia is a suite of decoder-only autoregressive language models all trained on public data seen in the exact same order and ranging in size from 70M to 12B parameters. The model architecture and hyperparameters largely follow GPT-3, with a few notable deviations based on recent advances in best practices for large scale language modeling.
Papers archive 2025-07-28
30 shown of 60, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
LexiMark: Robust Watermarking via Lexical Substitutions to Enhance Membership Verification of an LLM's Textual Training Data 17 Jun 2025 · 1 repository · arXiv:2506.14474
-
What Happens During the Loss Plateau? Understanding Abrupt Learning in Transformers 16 Jun 2025 · 1 repository · arXiv:2506.13688Syntology ran 0 of 2 samples · 2 unverified · 2 pointer-only (licence)
-
Stochastic Chameleons: Irrelevant Context Hallucinations Reveal Class-Based (Mis)Generalization in LLMs 28 May 2025 · 0 repositories · arXiv:2505.22630
-
Pretraining Language Models to Ponder in Continuous Space 27 May 2025 · 1 repository · arXiv:2505.20674Syntology ran 6 of 6 samples · 0 unverified · 3 pointer-only (licence)
-
Illusion or Algorithm? Investigating Memorization, Emergence, and Symbolic Processing in In-Context Learning 16 May 2025 · 1 repository · arXiv:2505.11004
-
Memorization or Interpolation ? Detecting LLM Memorization through Input Perturbation Analysis 5 May 2025 · 0 repositories · arXiv:2505.03019
-
An Empirical Study of the Role of Incompleteness and Ambiguity in Interactions with Large Language Models 23 Mar 2025 · 0 repositories · arXiv:2503.17936
-
I Predict Therefore I Am: Is Next Token Prediction Enough to Learn Human-Interpretable Concepts from Data? 12 Mar 2025 · 0 repositories · arXiv:2503.08980
-
PolyPythias: Stability and Outliers across Fifty Language Model Pre-Training Runs 12 Mar 2025 · 1 repository · arXiv:2503.09543
-
Interrogating LLM design under a fair learning doctrine 22 Feb 2025 · 0 repositories · arXiv:2502.16290
-
Revisiting Privacy, Utility, and Efficiency Trade-offs when Fine-Tuning Large Language Models 18 Feb 2025 · 0 repositories · arXiv:2502.13313
-
RoSTE: An Efficient Quantization-Aware Supervised Fine-Tuning Approach for Large Language Models 13 Feb 2025 · 0 repositories · arXiv:2502.09003Syntology ran 5 of 5 samples · 0 unverified · 5 pointer-only (licence)
-
MemHunter: Automated and Verifiable Memorization Detection at Dataset-scale in LLMs 10 Dec 2024 · 0 repositories · arXiv:2412.07261
-
Star-Agents: Automatic Data Optimization with LLM Agents for Instruction Tuning 21 Nov 2024 · 1 repository · arXiv:2411.14497
-
Explaining and Improving Contrastive Decoding by Extrapolating the Probabilities of a Huge and Hypothetical LM 3 Nov 2024 · 1 repository · arXiv:2411.01610
-
Efficient Training of Sparse Autoencoders for Large Language Models via Layer Groups 28 Oct 2024 · 0 repositories · arXiv:2410.21508
-
Relaxed Recursive Transformers: Effective Parameter Sharing with Layer-wise LoRA 28 Oct 2024 · 0 repositories · arXiv:2410.20672
-
Hallucination Detox: Sensitivity Dropout (SenD) for Large Language Model Training 20 Oct 2024 · 0 repositories · arXiv:2410.15460
-
Tending Towards Stability: Convergence Challenges in Small Language Models 15 Oct 2024 · 1 repository · arXiv:2410.11451
-
Context-Parametric Inversion: Why Instruction Finetuning Can Worsen Context Reliance 14 Oct 2024 · 0 repositories · arXiv:2410.10796
-
Large Language Model Evaluation via Matrix Nuclear-Norm 14 Oct 2024 · 1 repository · arXiv:2410.10672
-
Local and Global Decoding in Text Generation 14 Oct 2024 · 1 repository · arXiv:2410.10810Syntology ran 11 of 12 samples · 1 unverified · 12 pointer-only (licence)
-
Lightweight Deep Learning Framework for Accurate Particle Flow Energy Reconstruction 8 Oct 2024 · 1 repository · arXiv:2410.07250
-
Order of Magnitude Speedups for LLM Membership Inference 22 Sep 2024 · 0 repositories · arXiv:2409.14513
-
Generated Data with Fake Privacy: Hidden Dangers of Fine-tuning Large Language Models on Generated Data 12 Sep 2024 · 0 repositories · arXiv:2409.11423
-
Accelerating Large Language Model Pretraining via LFR Pedagogy: Learn, Focus, and Review 10 Sep 2024 · 0 repositories · arXiv:2409.06131
-
Demystifying Verbatim Memorization in Large Language Models 25 Jul 2024 · 1 repository · arXiv:2407.17817Syntology ran 6 of 6 samples · 0 unverified
-
Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data 20 Jul 2024 · 0 repositories · arXiv:2407.14985Syntology ran 11 of 15 samples · 4 unverified
-
MathCAMPS: Fine-grained Synthesis of Mathematical Problems From Human Curricula 1 Jul 2024 · 1 repository · arXiv:2407.00900Syntology ran 8 of 8 samples · 0 unverified
-
Evaluating n-Gram Novelty of Language Models Using Rusty-DAWG 18 Jun 2024 · 1 repository · arXiv:2406.13069
Tasks archive 2025-07-28
20 shown of 58 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections