Datasets › WenetSpeech
WenetSpeech
WenetSpeech is a multi-domain Mandarin corpus consisting of 10,000+ hours high-quality labeled speech, 2,400+ hours weakly labelled speech, and about 10,000 hours unlabeled speech, with 22,400+ hours in total. The authors collected the data from YouTube and Podcast, which covers a variety of speaking styles, scenarios, domains, topics, and noisy conditions. An optical character recognition (OCR) based method is introduced to generate the audio/text segmentation candidates for the YouTube data on its corresponding video captions.
Image source: https://github.com/wenet-e2e/wenetspeech
Benchmarks archive 2025-07-28
All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.
| First row (archive order) | Paper | Code | ||||
|---|---|---|---|---|---|---|
| Speech Recognition | WenetSpeech | Paraformer-large Character Error Rate (CER) 6.97 | FunASR: A Fundamental End-to-End Speech Recognition Toolkit | alibaba-damo-academy/FunASR | 8 | Compare |
Papers archive 2025-07-28
4 shown of 4 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 58. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.
| Date | Samples run Syntology | |||
|---|---|---|---|---|
| Zipformer: A faster and better encoder for automatic speech recognition | 1 | 1 | 17 Oct 2023 | not harvested |
| FunASR: A Fundamental End-to-End Speech Recognition Toolkit | 1 | 1 | 18 May 2023 | not harvested |
| 3M: Multi-loss, Multi-path and Multi-level Neural Networks for speech recognition | 1 | 3 | 7 Apr 2022 | ran 0 of 1 samples (1 unverified) |
| WenetSpeech: A 10000+ Hours Multi-domain Mandarin Corpus for Speech Recognition | 2 | 3 | 7 Oct 2021 | ran 0 of 8 samples (8 unverified) |
Dataset loaders archive 2025-07-28
1 loader as listed in the archive; links are outbound and not re-checked here.
Tasks archive 2025-07-28
License archive 2025-07-28
Modalities archive 2025-07-28
Languages archive 2025-07-28
Variants archive 2025-07-28
- WenetSpeech
1 variant name, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections