Datasets › TIMIT

TIMIT (TIMIT Acoustic-Phonetic Continuous Speech Corpus)

archive 2025-07-28

The TIMIT Acoustic-Phonetic Continuous Speech Corpus is a standard dataset used for evaluation of automatic speech recognition systems. It consists of recordings of 630 speakers of 8 dialects of American English each reading 10 phonetically-rich sentences. It also comes with the word and phone-level transcriptions of the speech.

Source: Improving neural networks by preventing co-adaptation of feature detectors Image Source: https://roboticrun.wordpress.com/2016/06/21/timit-introduction-the-official-doc/

Benchmarks archive 2025-07-28

All 6 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Speech Recognition TIMIT wav2vec 2.0 Percentage error 8.3 wav2vec 2.0: A Framework for Self-Supervised Learning of... huggingface/transformers +24 22 Compare
Lip Reading TCD-TIMIT corpus (mixed-speech) Lip2Wav WER 31.26 Learning Individual Speaking Styles for Accurate Lip to... Rudrabha/Lip2Wav 1 Compare
Speaker-Specific Lip to Speech Synthesis TCD-TIMIT corpus (mixed-speech) Lip2Wav ESTOI 36.5 Learning Individual Speaking Styles for Accurate Lip to... Rudrabha/Lip2Wav 1 Compare
Speech Enhancement TCD-TIMIT corpus (mixed-speech) Audio-Visual concat-ref PESQ 3.03 Face Landmark-based Speaker-Independent Audio-Visual... dr-pato/audio_visual_speech_enhancement 1 Compare
Speech Separation TCD-TIMIT corpus (mixed-speech) Audio-Visual concat-ref SDR 10.55 Face Landmark-based Speaker-Independent Audio-Visual... dr-pato/audio_visual_speech_enhancement 1 Compare
Speech Recognition DARPA TIMIT no rows — — 0 Compare

Papers archive 2025-07-28

14 shown of 14 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 31. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations 25 1 20 Jun 2020 ran 2 of 9 samples (7 unverified; 2 pointer-only for licence)
Learning Individual Speaking Styles for Accurate Lip to Speech Synthesis 1 2 17 May 2020 ran 1 of 9 samples (8 unverified)
vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations 3 1 12 Oct 2019 ran 0 of 3 samples (3 unverified)
Attention model for articulatory features detection 1 1 2 Jul 2019 not harvested
wav2vec: Unsupervised Pre-training for Speech Recognition 7 1 11 Apr 2019 ran 1 of 1 samples (0 unverified)
The PyTorch-Kaldi Speech Recognition Toolkit 11 8 19 Nov 2018 ran 5 of 6 samples (1 unverified; 6 pointer-only for licence)
Face Landmark-based Speaker-Independent Audio-Visual Speech Enhancement in Multi-Talker Environments 1 2 6 Nov 2018 not harvested
Quaternion Convolutional Neural Networks for End-to-End Automatic Speech Recognition 1 1 20 Jun 2018 not harvested
Long short-term memory and learning-to-learn in networks of spiking neurons 2 1 26 Mar 2018 ran 3 of 5 samples (2 unverified; 5 pointer-only for licence)
Light Gated Recurrent Units for Speech Recognition 1 2 26 Mar 2018 not harvested
Online and Linear-Time Attention by Enforcing Monotonic Alignments 2 1 3 Apr 2017 not harvested
Segmental Recurrent Neural Networks for End-to-end Speech Recognition 0 1 1 Mar 2016 not harvested
Attention-Based Models for Speech Recognition 14 1 24 Jun 2015 not harvested
Speech Recognition with Deep Recurrent Neural Networks 5 1 22 Mar 2013 ran 0 of 3 samples (3 unverified)

Dataset loaders archive 2025-07-28

3 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • timit PER
  • TIMIT
  • TCD-TIMIT corpus (mixed-speech)
  • DARPA TIMIT

4 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections