Papers › Language Models are Few-Shot Learners
Language Models are Few-Shot Learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, Dario Amodei
Recent work has demonstrated substantial gains on many NLP tasks and benchmarks by pre-training on a large corpus of text followed by fine-tuning on a specific task. While typically task-agnostic in architecture, this method still requires task-specific fine-tuning datasets of thousands or tens of thousands of examples. By contrast, humans can generally perform a new language task from only a few examples or from simple instructions - something which current NLP systems still largely struggle to do. Here we show that scaling up language models greatly improves task-agnostic, few-shot performance, sometimes even reaching competitiveness with prior state-of-the-art fine-tuning approaches. Specifically, we train GPT-3, an autoregressive language model with 175 billion parameters, 10x more than any previous non-sparse language model, and test its performance in the few-shot setting. For all tasks, GPT-3 is applied without any gradient updates or fine-tuning, with tasks and few-shot demonstrations specified purely via text interaction with the model. GPT-3 achieves strong performance on many NLP datasets, including translation, question-answering, and cloze tasks, as well as several tasks that require on-the-fly reasoning or domain adaptation, such as unscrambling words, using a novel word in a sentence, or performing 3-digit arithmetic. At the same time, we also identify some datasets where GPT-3's few-shot learning still struggles, as well as some datasets where GPT-3 faces methodological issues related to training on large web corpora. Finally, we find that GPT-3 can generate samples of news articles which human evaluators have difficulty distinguishing from articles written by humans. We discuss broader societal impacts of this finding and of GPT-3 in general.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
For agents, Syntology's MCP tool lists every function and class Syntology harvested from this paper and whether it ran (how to connect): get_harvested_code_for_paper(arxiv_id="2005.14165")
Code
Syntology Ran 15 of 65 code samples harvested from 17 repositories linked to this paper; 50 have no recorded run. Of those that ran: 1 ran · honoured contract; 1 ran · violated contract; 4 ran · our draft was wrong; 1 ran · fixture could not drive it; 8 ran with no contract checked.
By repository: community (archive-listed): 61 samples from 17 repositories, 15 ran; 4 identical to code first harvested elsewhere. The run record, sample by sample. “Ran” means executed on a synthesized input, not that the code is correct or reproduces the paper.
67 repositories listed; official and paper-mentioned ones first.
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
65 samples harvested; 15 ran; 1 honoured the contract we drafted; 50 have no recorded run. Read from Syntology's graph 2026-09-24; that is when this build read the record, not when the samples ran.
Licence: 4 of the 65 samples are pointer only, meaning Syntology does not serve that copy's text. This page shows no code text for any sample; each one links to its file in the repository.
Harvested from 17 repositories linked to this paper, official or community; each sample names its own and says which. Some samples are identical code Syntology first harvested from another repository; for those, this paper's copy is not located and its licence is not recorded. “Ran” means the sample executed on a synthesized input. It does not mean the output is correct, and nothing here reproduces the paper's results. “Honoured” and “violated” refer to a contract Syntology drafted from the code itself; “our draft was wrong” and “fixture could not drive it” are failures of Syntology's instrument, not of the code.
Each sample ends with its code_sha256, Syntology's identity for that exact code. An agent fetches the stored sample with Syntology's MCP tool get_code(code_sha256="…") (how to connect); click an identity to copy that call.
Repository labels, per sample. official repository: The archive marks this repository official for the paper. named in the paper: The archive records that the paper mentions this repository; it is not marked official. community (archive-listed): In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper. found in paper text by Syntology: Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted. community: Not in the archive's code links for this paper; a community repository Syntology harvested. Samples from a repository marked official are listed first. Licence labels name the repository's licence as recorded at harvest. “Pointer only” means Syntology does not serve that copy's text, for one of four reasons: no licence file was found; the licence was not identified; the licence is recorded as permissive but that copy's record is not marked cleared; or the licence is outside the permissive list Syntology serves text under (MIT, Apache-2.0, BSD and similar). Some licences outside that list permit redistribution, such as WTFPL, and GPL-3.0 under its conditions; they are simply not on the list. Hover a licence label for the reason. File links open the file on GitHub at the default branch, which may have changed since the harvest.
f0fc0ac7590a2d39 · report
57e732507bad40db · report
53354f082d9addc6 · report
00f5e7d9ee6636ba · report
5b315604fb608f3b · report
0ee5d4d20cf990bd · report
20a7cc804eb22661 · report
ea06eaae4fc1eaf0 · report
cc82794b6ec72130 · report
400ea4f64ccce5ef · report
e60f912106425ef4 · report
809bc9208dd99dce · report
de43c7cdbf869350 · report
c44830a7722ff5f7 · report
dfed6e9291316231 · report
3bb8fcfcb4c12121 · report
e66a197775436632 · report
7ca8b4cd5f558e02 · report
9dd643ee2669464c · report
904fac6af5735d5a · report
64087979238e1d41 · report
451bc0edacfc317d · report
3055a2ccdae837c7 · report
3e6ff0c8581ee623 · report
ffcac685bcb61388 · report
a20ac4af0aec2beb · report
ecbbe3133a0f68b3 · report
5307f83d52fe15d1 · report
03edfbaef45e2c28 · report
3866cbd2c9f5fd58 · report
34f65890e3e58ec4 · report
87cfebcf19ab803b · report
c7271f2f2ce3c4fa · report
c56b548854fce3e6 · report
f6825b4eb5fe78c8 · report
306b51e6dcfb30ef · report
b1e0531f8c17d5d6 · report
8e3012dd13fbb68a · report
5decf8cec7692d54 · report
89b22f3da55a855c · report
0db35c379a280bd8 · report
0fb487fd2495f33f · report
26519d22d8af6d48 · report
72f5b387b64e9d9b · report
3e9d38ecacd03fa7 · report
2e6ad295f92e3d9d · report
deba15e38b692b23 · report
4ff6d4e326d3c0db · report
2ee40684ab58dcb9 · report
fdde66f5d1999801 · report
bd67f40255aa71d3 · report
7cacc3b113e10007 · report
63b45afd720a808e · report
99b9d9d7a4d6ca29 · report
9d752127437e6b8b · report
c7a7a9dca42a1347 · report
2fe668161976f857 · report
5036094f3af0dc56 · report
cabc905de3f72170 · report
11b1822611370e59 · report
8cb749113ad26cc3 · report
e1d89d7579eaaeec · report
7753000f9482cac5 · report
77347b697708c203 · report
d09226e56c2085ea · report
Tasks
Results from the paper archive 2025-07-28
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Methods
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections