Papers › Robust Speech Recognition via Large-Scale Weak Supervision
Robust Speech Recognition via Large-Scale Weak Supervision
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, Ilya Sutskever
We study the capabilities of speech processing systems trained simply to predict large amounts of transcripts of audio on the internet. When scaled to 680,000 hours of multilingual and multitask supervision, the resulting models generalize well to standard benchmarks and are often competitive with prior fully supervised results but in a zero-shot transfer setting without the need for any fine-tuning. When compared to humans, the models approach their accuracy and robustness. We are releasing models and inference code to serve as a foundation for further work on robust speech processing.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
For agents, Syntology's MCP tool lists every function and class Syntology harvested from this paper and whether it ran (how to connect): get_harvested_code_for_paper(arxiv_id="2212.04356")
Code
Syntology Ran 5 of 59 code samples harvested from 9 repositories linked to this paper; 54 have no recorded run. Of those that ran: 2 ran · honoured contract; 1 ran · our draft was wrong; 1 ran · fixture could not drive it; 1 ran with no contract checked.
By repository: community (archive-listed): 58 samples from 9 repositories, 4 ran; 1 identical to code first harvested elsewhere. The run record, sample by sample. “Ran” means executed on a synthesized input, not that the code is correct or reproduces the paper.
15 repositories listed; official and paper-mentioned ones first.
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
59 samples harvested; 5 ran; 2 honoured the contract we drafted; 54 have no recorded run. Read from Syntology's graph 2026-09-24; that is when this build read the record, not when the samples ran.
Licence: 18 of the 59 samples are pointer only, meaning Syntology does not serve that copy's text. This page shows no code text for any sample; each one links to its file in the repository.
Harvested from 9 repositories linked to this paper, official or community; each sample names its own and says which. Some samples are identical code Syntology first harvested from another repository; for those, this paper's copy is not located and its licence is not recorded. “Ran” means the sample executed on a synthesized input. It does not mean the output is correct, and nothing here reproduces the paper's results. “Honoured” and “violated” refer to a contract Syntology drafted from the code itself; “our draft was wrong” and “fixture could not drive it” are failures of Syntology's instrument, not of the code.
Each sample ends with its code_sha256, Syntology's identity for that exact code. An agent fetches the stored sample with Syntology's MCP tool get_code(code_sha256="…") (how to connect); click an identity to copy that call.
Repository labels, per sample. official repository: The archive marks this repository official for the paper. named in the paper: The archive records that the paper mentions this repository; it is not marked official. community (archive-listed): In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper. found in paper text by Syntology: Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted. community: Not in the archive's code links for this paper; a community repository Syntology harvested. Samples from a repository marked official are listed first. Licence labels name the repository's licence as recorded at harvest. “Pointer only” means Syntology does not serve that copy's text, for one of four reasons: no licence file was found; the licence was not identified; the licence is recorded as permissive but that copy's record is not marked cleared; or the licence is outside the permissive list Syntology serves text under (MIT, Apache-2.0, BSD and similar). Some licences outside that list permit redistribution, such as WTFPL, and GPL-3.0 under its conditions; they are simply not on the list. Hover a licence label for the reason. File links open the file on GitHub at the default branch, which may have changed since the harvest.
2000f34f95c683d6 · report
6e0808fc4828513a · report
8a8edc42e5e25cea · report
e529c9641c178fe0 · report
6ecbcf4f3d00f7ed · report
5969d689c8a84ef2 · report
69124b03f4b749bd · report
41c654294ba9ec56 · report
3787ac531561b6c6 · report
9f16788e5da13dfa · report
acb6a55d43c7f465 · report
59d410aa19bdef96 · report
004eb0067bd43754 · report
9140147b0fd2c440 · report
fa47326ca5e0886b · report
c4c3b54ad53b1fce · report
5632d6609129dfa3 · report
8a64b592c82c0d10 · report
b41e987fee8a92ff · report
76330ed965e60af5 · report
87c9b24deaf27799 · report
9649e15864c56627 · report
fef7828e835c938d · report
63302892207b7f5b · report
b865558252e1566a · report
ae28fa768d63729b · report
ad63708b6da385d5 · report
b3916c1b2edcf996 · report
ee3c8d2a605ac8bc · report
a5b8c22a3ac7b238 · report
d1162925e2d53784 · report
f2f52f6a90834243 · report
9720e570a35f6d55 · report
a799a46446fafd99 · report
a972a5df2771f7a5 · report
34dd5baae0e5830e · report
ba1be9f251ecd80a · report
a1f5678203ff47e8 · report
c45f3b632ec36534 · report
6f669ba80cbbf823 · report
edf924293b4b2375 · report
40e421c01b760f1c · report
c4ee356de1bd8c77 · report
e56996b20b9c19cb · report
9f13de1c23e4003d · report
87bdf3970749dc11 · report
d35e5ea9b14ce05e · report
2206d84b6900ab3e · report
cef27f679a755971 · report
e65744d3337badd5 · report
bdbf18d0ee9006f7 · report
407b4297398ff29d · report
226350299950338b · report
176b6ba6c2c2940c · report
2ef9c9c7ee284aaf · report
3d1de040cc006f6a · report
738c6e74fd16381d · report
e5a9a9806511efb5 · report
0b5f90e8856ad7af · report
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Speech Recognition | Common Voice English | Whisper (Large v2) | Word Error Rate (WER) | 9.4% | #2 of 2 | Archive leaderboard | report |
| Speech Recognition | Common Voice French | Whisper (Large v2) | Test WER | 13.9% | #8 of 8 | Archive leaderboard | report |
| Speech Recognition | Common Voice German | Whisper (Large v2) | Test WER | 6.4% | #7 of 14 | Archive leaderboard | report |
| Speech Recognition | Common Voice Italian | Whisper (Large v2) | Test WER | 7.1% | #1 of 2 | Archive leaderboard | report |
| Speech Recognition | Common Voice Japanese | Whisper (Large v2) | Test WER | 9.1% | #1 of 1 | Archive leaderboard | report |
| Speech Recognition | Common Voice Russian | Whisper (Large v2) | Test WER | 7.1% | #1 of 1 | Archive leaderboard | report |
| Speech Recognition | Common Voice Spanish | Whisper (Large v2) | Test WER | 5.6% | #2 of 8 | Archive leaderboard | report |
| Speech-to-Speech Translation | FLEURS X-eng | WhisperV2 | ASR-BLEU | 23.5 | #6 of 7 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections