Methods › Natural Language Processing › Language Models › OPT

OPT

285 papers tagged archive 2025-07-28

Introduced by Susan Zhang et al. in OPT: Open Pre-trained Transformer Language Models

archive 2025-07-28 Description, source and code snippet are the archive's method entry.

OPT is a suite of decoder-only pre-trained transformers ranging from 125M to 175B parameters. The model uses an AdamW optimizer and weight decay of 0.1. It follows a linear learning rate schedule, warming up from 0 to the maximum learning rate over the first 2000 steps in OPT-175B, or over 375M tokens in the smaller models, and decaying down to 10% of the maximum LR over 300B tokens. The batch sizes range from 0.5M to 4M depending on the model size and is kept constant throughout the course of training.

PaperSource

Papers archive 2025-07-28

30 shown of 285, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.

Tasks archive 2025-07-28

20 shown of 247 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.

TaskPapers
Language Modelling41
Language Modeling28
Quantization24
GPU22
Large Language Model17
In-Context Learning14
Question Answering14
Text Generation12
Model Compression8
Retrieval8
Fairness7
Transfer Learning7
Translation7
Attribute6
Benchmarking6
Object Detection6
Scheduling6
Sentence6
object-detection6
Computational Efficiency5

Usage over time archive 2025-07-28

Papers per year tagged with OPT: 2022 to 2025, peak 104 104 0 2022: 49 papers 2022 2023: 104 papers 2023 2024: 104 papers 2024 2025: 28 papers 2025
Papers per year the archive tags with this method, by the paper's archive date (285 dated). Bars are counts, not a trend claim.

Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).

Categories archive 2025-07-28

Language Models

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections