Methods › Natural Language Processing › Language Models › OPT
OPT
Introduced by Susan Zhang et al. in OPT: Open Pre-trained Transformer Language Models
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
OPT is a suite of decoder-only pre-trained transformers ranging from 125M to 175B parameters. The model uses an AdamW optimizer and weight decay of 0.1. It follows a linear learning rate schedule, warming up from 0 to the maximum learning rate over the first 2000 steps in OPT-175B, or over 375M tokens in the smaller models, and decaying down to 10% of the maximum LR over 300B tokens. The batch sizes range from 0.5M to 4M depending on the model size and is kept constant throughout the course of training.
Papers archive 2025-07-28
30 shown of 285, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
Incentivizing High-quality Participation From Federated Learning Agents 20 Jun 2025 · 0 repositories · arXiv:2506.16731
-
Attribution-guided Pruning for Compression, Circuit Discovery, and Targeted Correction in LLMs 16 Jun 2025 · 1 repository · arXiv:2506.13727
-
Mixture of Weight-shared Heterogeneous Group Attention Experts for Dynamic Token-wise KV Optimization 16 Jun 2025 · 0 repositories · arXiv:2506.13541
-
TensorSLM: Energy-efficient Embedding Compression of Sub-billion Parameter Language Models on Low-end Devices 16 Jun 2025 · 0 repositories · arXiv:2506.13514
-
FZOO: Fast Zeroth-Order Optimizer for Fine-Tuning Large Language Models towards Adam-Scale Speed 10 Jun 2025 · 0 repositories · arXiv:2506.09034
-
Conservative Bias in Large Language Models: Measuring Relation Predictions 9 Jun 2025 · 0 repositories · arXiv:2506.08120
-
NQKV: A KV Cache Quantization Scheme Based on Normal Distribution Characteristics 22 May 2025 · 0 repositories · arXiv:2505.16210
-
Polar Sparsity: High Throughput Batched LLM Inferencing with Scalable Contextual Sparsity 20 May 2025 · 1 repository · arXiv:2505.14884Syntology ran 3 of 3 samples · 0 unverified
-
Optimizing Energy Consumption in Stochastic Production Systems: Using a Simulation-Based Approach for Stopping Policy 14 May 2025 · 0 repositories · arXiv:2505.11536
-
Anticipating Gaming to Incentivize Improvement: Guiding Agents in (Fair) Strategic Classification 8 May 2025 · 0 repositories · arXiv:2505.05594
-
SPAP: Structured Pruning via Alternating Optimization and Penalty Methods 6 May 2025 · 0 repositories · arXiv:2505.03373
-
Efficient Shapley Value-based Non-Uniform Pruning of Large Language Models 3 May 2025 · 0 repositories · arXiv:2505.01731
-
UniDetox: Universal Detoxification of Large Language Models via Dataset Distillation 29 Apr 2025 · 1 repository · arXiv:2504.20500Syntology ran 1 of 1 samples · 0 unverified
-
PIN-WM: Learning Physics-INformed World Models for Non-Prehensile Manipulation 23 Apr 2025 · 0 repositories · arXiv:2504.16693
-
Transferable text data distillation by trajectory matching 14 Apr 2025 · 0 repositories · arXiv:2504.09818
-
Maximum Redundancy Pruning: A Principle-Driven Layerwise Sparsity Allocation for LLMs 24 Mar 2025 · 0 repositories · arXiv:2503.18377
-
MuMA: 3D PBR Texturing via Multi-Channel Multi-View Generation and Agentic Post-Processing 24 Mar 2025 · 0 repositories · arXiv:2503.18461
-
How does Bike Absence Influence Mode Shifts Among Dockless Bike-Sharing Users? Evidence From Nanjing, China 18 Mar 2025 · 0 repositories · arXiv:2503.14265
-
Large language models in finance : what is financial sentiment? 5 Mar 2025 · 0 repositories · arXiv:2503.03612
-
Probe Pruning: Accelerating LLMs through Dynamic Pruning via Model-Probing 21 Feb 2025 · 1 repository · arXiv:2502.15618Syntology ran 2 of 3 samples · 1 unverified · 1 pointer-only (licence)
-
Enhancing Cross-Tokenizer Knowledge Distillation with Contextual Dynamical Mapping 16 Feb 2025 · 1 repository · arXiv:2502.11104
-
The Odyssey of the Fittest: Can Agents Survive and Still Be Good? 8 Feb 2025 · 1 repository · arXiv:2502.05442
-
GREEN-CODE: Learning to Optimize Energy Efficiency in LLM-based Code Generation 19 Jan 2025 · 1 repository · arXiv:2501.11006
-
FASP: Fast and Accurate Structured Pruning of Large Language Models 16 Jan 2025 · 0 repositories · arXiv:2501.09412
-
Multilingual LLMs Struggle to Link Orthography and Semantics in Bilingual Word Processing 15 Jan 2025 · 1 repository · arXiv:2501.09127
-
Has an AI model been trained on your images? 11 Jan 2025 · 0 repositories · arXiv:2501.06399
-
Explore Activation Sparsity in Recurrent LLMs for Energy-Efficient Neuromorphic Computing 9 Jan 2025 · 1 repository · arXiv:2501.16337
-
From Sparse Signal to Smooth Motion: Real-Time Motion Generation with Rolling Prediction Models 1 Jan 2025 · 0 repositories
-
Sentiment trading with large language models 26 Dec 2024 · 0 repositories · arXiv:2412.19245
-
RAGONITE: Iterative Retrieval on Induced Databases and Verbalized RDF for Conversational QA over KGs with RAG 23 Dec 2024 · 0 repositories · arXiv:2412.17690
Tasks archive 2025-07-28
20 shown of 247 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections