Datasets › WenetSpeech

WenetSpeech

Introduced by BinBin Zhang et al. in WenetSpeech: A 10000+ Hours Multi-domain Mandarin Corpus for Speech Recognition7 Oct 2021 archive 2025-07-28

WenetSpeech is a multi-domain Mandarin corpus consisting of 10,000+ hours high-quality labeled speech, 2,400+ hours weakly labelled speech, and about 10,000 hours unlabeled speech, with 22,400+ hours in total. The authors collected the data from YouTube and Podcast, which covers a variety of speaking styles, scenarios, domains, topics, and noisy conditions. An optical character recognition (OCR) based method is introduced to generate the audio/text segmentation candidates for the YouTube data on its corresponding video captions.

Image source: https://github.com/wenet-e2e/wenetspeech

Benchmarks archive 2025-07-28

All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Speech Recognition WenetSpeech Paraformer-large Character Error Rate (CER) 6.97 FunASR: A Fundamental End-to-End Speech Recognition Toolkit alibaba-damo-academy/FunASR 8 Compare

Papers archive 2025-07-28

4 shown of 4 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 58. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Zipformer: A faster and better encoder for automatic speech recognition 1 1 17 Oct 2023 not harvested
FunASR: A Fundamental End-to-End Speech Recognition Toolkit 1 1 18 May 2023 not harvested
3M: Multi-loss, Multi-path and Multi-level Neural Networks for speech recognition 1 3 7 Apr 2022 ran 0 of 1 samples (1 unverified)
WenetSpeech: A 10000+ Hours Multi-domain Mandarin Corpus for Speech Recognition 2 3 7 Oct 2021 ran 0 of 8 samples (8 unverified)

Dataset loaders archive 2025-07-28

1 loader as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

CC BY 4.0

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • WenetSpeech

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections