Datasets › 2000 HUB5 English

2000 HUB5 English

archive 2025-07-28

2000 HUB5 English Evaluation Transcripts was developed by the Linguistic Data Consortium (LDC) and consists of transcripts of 40 English telephone conversations used in the 2000 HUB5 evaluation sponsored by NIST (National Institute of Standards and Technology).

The Hub5 evaluation series focused on conversational speech over the telephone with the particular task of transcribing conversational speech into text. Its goals were to explore promising new areas in the recognition of conversational speech, to develop advanced technology incorporating those ideas and to measure the performance of new technology.

Benchmarks archive 2025-07-28

All 2 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Speech Recognition Hub5'00 SwitchBoard LAS + SpecAugment (with LM, Switchboard mild policy) SwitchBoard 6.8 SpecAugment: A Simple Data Augmentation Method for... mozilla/DeepSpeech +29 5 Compare
Language Modelling 2000 HUB5 English MMLU 10-stage average accuracy 10 Spirit LM: Interleaved Spoken and Written Language Model facebookresearch/spiritlm 1 Compare

Papers archive 2025-07-28

5 shown of 5 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 33. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Spirit LM: Interleaved Spoken and Written Language Model 1 1 8 Feb 2024 ran 1 of 1 samples (0 unverified; 1 pointer-only for licence)
CAT: A CTC-CRF based ASR Toolkit Bridging the Hybrid and the End-to-end Approaches towards Data Efficiency and Low Latency 1 1 27 May 2020 not harvested
Espresso: A Fast End-to-end Neural Speech Recognition Toolkit 1 1 18 Sep 2019 ran 1 of 1 samples (0 unverified; 1 pointer-only for licence)
SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition 30 2 18 Apr 2019 ran 1 of 18 samples (17 unverified)
Jasper: An End-to-End Convolutional Neural Acoustic Model 10 1 5 Apr 2019 not harvested

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • Hub5'00 SwitchBoard
  • 2000 HUB5 English
  • SWBD Eval2000

3 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections