Browse State-of-the-Art › Speech
Speech
Benchmarks are leaderboard tables with at least one row whose task is in this area, counted by the task's area and not by the archive's per-table category tag, which tags 409 tables with Speech; datasets are those the archive tags with a task in this area; papers with code are catalogue papers tagged with a task in this area that list at least one repository. Task images are not shown (the archive's image host no longer serves them).
Parent tasks
21 tasks in Speech have sub-tasks in the archive's task tree, most benchmarks first, then most papers with code. Each section shows up to 5 sub-tasks; the task page lists them all.
- Speech Recognition (11)
- Text Generation (25)
- Speech Separation (1)
- Speech Enhancement (4)
- Speech Emotion Recognition (4)
- Speaker Verification (3)
- Keyword Spotting (2)
- Automatic Speech Recognition (ASR) (1)
- Emotion Recognition (12)
- DeepFake Detection (5)
- Text-To-Speech Synthesis (2)
- Speech Synthesis (15)
- Spoken Language Understanding (2)
- Audio-Visual Speech Recognition (1)
- Audio Generation (4)
- Visual Speech Recognition (1)
- Chatbot (1)
- 1 Image, 2*2 Stitchi (12)
- Dialogue (21)
- Speaker Separation (1)
- Pronunciation Assessment (3)
Speech Recognition
65 benchmarks · 1,373 papers with codeAutomatic Speech Recognition (ASR)
9 benchmarks · 622 papers with code
Automatic Lyrics Transcription
5 benchmarks · 9 papers with code
Accented Speech Recognition
4 benchmarks · 6 papers with code
Visual Speech Recognition
2 benchmarks · 62 papers with code
Distant Speech Recognition
2 benchmarks · 10 papers with code
5 shown of 11 sub-tasks (1 filed under another area). All sub-tasks of Speech Recognition →
Text Generation
20 benchmarks · 2,047 papers with codeData-to-Text Generation
26 benchmarks · 112 papers with code
Dialogue Generation
13 benchmarks · 265 papers with code
Table-to-Text Generation
8 benchmarks · 43 papers with code
Code Documentation Generation
7 benchmarks · 7 papers with code
Multi-Document Summarization
5 benchmarks · 113 papers with code
5 shown of 25 sub-tasks (10 filed under another area). All sub-tasks of Text Generation →
Speech Separation
19 benchmarks · 120 papers with codeSpeech Extraction
1 benchmark · 10 papers with code
1 shown of 1 sub-task.
Speech Enhancement
17 benchmarks · 280 papers with codeBandwidth Extension
6 benchmarks · 18 papers with code
Speech Dereverberation
5 benchmarks · 20 papers with code
Packet Loss Concealment
0 benchmarks · 4 papers with code
Speech Intelligibility Evaluation
0 benchmarks · 0 papers with code
4 shown of 4 sub-tasks.
Speech Emotion Recognition
16 benchmarks · 139 papers with codeVocal Bursts Intensity Prediction
1 benchmark · 775 papers with code
Vocal Bursts Valence Prediction
1 benchmark · 265 papers with code
Vocal Bursts Type Prediction
1 benchmark · 155 papers with code
Cultural Vocal Bursts Intensity Prediction
1 benchmark · 93 papers with code
4 shown of 4 sub-tasks.
Speaker Verification
12 benchmarks · 200 papers with codeAudio Deepfake Detection
2 benchmarks · 35 papers with code
Text-Independent Speaker Verification
0 benchmarks · 19 papers with code
Text-Dependent Speaker Verification
0 benchmarks · 2 papers with code
3 shown of 3 sub-tasks.
Keyword Spotting
10 benchmarks · 113 papers with codeVisual Keyword Spotting
3 benchmarks · 4 papers with code
Small-Footprint Keyword Spotting
0 benchmarks · 9 papers with code
2 shown of 2 sub-tasks.
Automatic Speech Recognition (ASR)
9 benchmarks · 622 papers with codeAutomatic Phoneme Recognition
6 benchmarks · 1 paper with code
1 shown of 1 sub-task.
Emotion Recognition
9 benchmarks · 614 papers with codeSpeech Emotion Recognition
16 benchmarks · 139 papers with code
Emotion Recognition in Conversation
16 benchmarks · 83 papers with code
Multimodal Emotion Recognition
7 benchmarks · 80 papers with code
Emotion Recognition in Context
4 benchmarks · 5 papers with code
EEG Emotion Recognition
3 benchmarks · 14 papers with code
5 shown of 12 sub-tasks (1 filed under another area). All sub-tasks of Emotion Recognition →
DeepFake Detection
8 benchmarks · 237 papers with codeAudio Deepfake Detection
2 benchmarks · 35 papers with code
Multimodal Forgery Detection
1 benchmark · 1 paper with code
Synthetic Speech Detection
0 benchmarks · 12 papers with code
diffusion-generated faces detection
0 benchmarks · 1 paper with code
Human Detection of Deepfakes
0 benchmarks · 1 paper with code
5 shown of 5 sub-tasks.
Text-To-Speech Synthesis
6 benchmarks · 104 papers with codeProsody Prediction
1 benchmark · 4 papers with code
Zero-Shot Multi-Speaker TTS
0 benchmarks · 3 papers with code
2 shown of 2 sub-tasks.
Speech Synthesis
5 benchmarks · 366 papers with codeSpeech Synthesis - Gujarati
2 benchmarks · 2 papers with code
Speech Synthesis - Assamese
1 benchmark · 1 paper with code
Speech Synthesis - Bengali
1 benchmark · 1 paper with code
Speech Synthesis - Bodo
1 benchmark · 1 paper with code
Speech Synthesis - Hindi
1 benchmark · 1 paper with code
5 shown of 15 sub-tasks. All sub-tasks of Speech Synthesis →
Spoken Language Understanding
5 benchmarks · 135 papers with codeSpoken language identification
12 benchmarks · 13 papers with code
Speech Tokenization
0 benchmarks · 9 papers with code
2 shown of 2 sub-tasks.
Audio-Visual Speech Recognition
4 benchmarks · 42 papers with codeText to Speech
2 benchmarks · 399 papers with code
1 shown of 1 sub-task.
Audio Generation
3 benchmarks · 124 papers with codeAudio Super-Resolution
4 benchmarks · 16 papers with code
Video-to-Sound Generation
1 benchmark · 6 papers with code
Voice Cloning
0 benchmarks · 36 papers with code
Room Impulse Response (RIR)
0 benchmarks · 18 papers with code
4 shown of 4 sub-tasks (1 filed under another area).
Visual Speech Recognition
2 benchmarks · 62 papers with codeLip to Speech Synthesis
1 benchmark · 6 papers with code
1 shown of 1 sub-task (1 filed under another area).
Chatbot
1 benchmark · 269 papers with codeDialogue Generation
13 benchmarks · 265 papers with code
1 shown of 1 sub-task (1 filed under another area).
1 Image, 2*2 Stitchi
1 benchmark · 3 papers with codePose Estimation
31 benchmarks · 1,679 papers with code
Text-to-Image Generation
17 benchmarks · 546 papers with code
Image Deblurring
9 benchmarks · 167 papers with code
Virtual Try-on
9 benchmarks · 114 papers with code
Style Transfer
3 benchmarks · 759 papers with code
5 shown of 12 sub-tasks (9 filed under another area). All sub-tasks of 1 Image, 2*2 Stitchi →
Dialogue
1 benchmark · 1 paper with codeDialogue Generation
13 benchmarks · 265 papers with code
Visual Dialog
8 benchmarks · 56 papers with code
Dialogue State Tracking
7 benchmarks · 138 papers with code
Dialogue Act Classification
5 benchmarks · 23 papers with code
Task-Oriented Dialogue Systems
4 benchmarks · 131 papers with code
5 shown of 21 sub-tasks (6 filed under another area). All sub-tasks of Dialogue →
Speaker Separation
0 benchmarks · 13 papers with codeMulti-Speaker Source Separation
0 benchmarks · 6 papers with code
1 shown of 1 sub-task.
Pronunciation Assessment
0 benchmarks · 0 papers with codePhone-level pronunciation scoring
1 benchmark · 5 papers with code
Utterance-level pronounciation scoring
1 benchmark · 2 papers with code
Word-level pronunciation scoring
1 benchmark · 2 papers with code
3 shown of 3 sub-tasks.
Tasks with no parent task
27 tasks in Speech sit at the top of the archive's task tree with no sub-tasks of their own, most benchmarks first, then most papers with code.
Speaker Diarization
12 benchmarks · 93 papers with code
Speaker Identification
4 benchmarks · 74 papers with code
Speech-to-Speech Translation
3 benchmarks · 39 papers with code
Speech Denoising
2 benchmarks · 32 papers with code
Arabic Text Diacritization
2 benchmarks · 7 papers with code
Speaker Recognition
1 benchmark · 102 papers with code
Spoken Command Recognition
1 benchmark · 5 papers with code
Voice Query Recognition
1 benchmark · 2 papers with code
Speech Representation Learning
0 benchmarks · 48 papers with code
Singing Voice Synthesis
0 benchmarks · 28 papers with code
X-ray Classification
0 benchmarks · 27 papers with code
Spoken Dialogue Systems
0 benchmarks · 26 papers with code
Acoustic echo cancellation
0 benchmarks · 12 papers with code
Acoustic Modelling
0 benchmarks · 11 papers with code
Speaker anonymization
0 benchmarks · 10 papers with code
Unsupervised Speech Recognition
0 benchmarks · 7 papers with code
Text-Independent Speaker Recognition
0 benchmarks · 6 papers with code
Voice Similarity
0 benchmarks · 5 papers with code
Music Genre Transfer
0 benchmarks · 4 papers with code
Silent Speech Recognition
0 benchmarks · 3 papers with code
Speaker Profiling
0 benchmarks · 3 papers with code
Acoustic Question Answering
0 benchmarks · 2 papers with code
Manner Of Articulation Detection
0 benchmarks · 2 papers with code
Speech Interruption Detection
0 benchmarks · 1 paper with code
Speech-to-Gesture Translation
0 benchmarks · 1 paper with code
When should a hot water tank be replaced?
0 benchmarks · 1 paper with code
Speaking Style Synthesis
0 benchmarks · 0 papers with code
3 tasks in Speech are filed only under a parent task from another area and are not listed on this page; the parent's task page carries them.
Task tree and counts are the archive's, frozen 2025-07-28 archive 2025-07-28. Nothing here is re-ranked.