Browse State-of-the-Art › Audio
Audio
Benchmarks are leaderboard tables with at least one row whose task is in this area, counted by the task's area and not by the archive's per-table category tag, which tags 488 tables with Audio; datasets are those the archive tags with a task in this area; papers with code are catalogue papers tagged with a task in this area that list at least one repository. Task images are not shown (the archive's image host no longer serves them).
Parent tasks
33 tasks in Audio have sub-tasks in the archive's task tree, most benchmarks first, then most papers with code. Each section shows up to 5 sub-tasks; the task page lists them all.
- Speech Recognition (11)
- Classification (24)
- Few-Shot Learning (12)
- Audio Classification (4)
- Speech Enhancement (4)
- 2D Semantic Segmentation (17)
- Emotion Recognition (12)
- DeepFake Detection (5)
- Language Identification (2)
- Text-To-Speech Synthesis (2)
- Speech Synthesis (15)
- Text to Audio Retrieval (1)
- Accented Speech Recognition (1)
- Audio Generation (4)
- Environmental Sound Classification (1)
- 10-shot image generation (21)
- Target Sound Extraction (1)
- Online Beat Tracking (1)
- Audio captioning (2)
- Audio Source Separation (3)
- Music Generation (3)
- 3D Human Pose Tracking (1)
- 1 Image, 2*2 Stitchi (12)
- Directional Hearing (1)
- Audio Signal Processing (3)
- Instance Search (1)
- 3D Human Dynamics (1)
- Audio Effects Modeling (2)
- Audio Signal Recognition (1)
- 2D Classification (27)
- Bird Classification (2)
- Hearing Aid and device processing (2)
- Signal Processing (1)
Speech Recognition
65 benchmarks · 1,373 papers with codeAutomatic Speech Recognition (ASR)
9 benchmarks · 622 papers with code
Automatic Lyrics Transcription
5 benchmarks · 9 papers with code
Accented Speech Recognition
4 benchmarks · 6 papers with code
Visual Speech Recognition
2 benchmarks · 62 papers with code
Distant Speech Recognition
2 benchmarks · 10 papers with code
5 shown of 11 sub-tasks (2 filed under another area). All sub-tasks of Speech Recognition →
Classification
58 benchmarks · 3,778 papers with codeGraph Classification
73 benchmarks · 483 papers with code
Text Classification
68 benchmarks · 1,308 papers with code
Audio Classification
22 benchmarks · 183 papers with code
Medical Image Classification
11 benchmarks · 183 papers with code
Multi-class Classification
5 benchmarks · 289 papers with code
5 shown of 24 sub-tasks (5 filed under another area). All sub-tasks of Classification →
Few-Shot Learning
27 benchmarks · 1,297 papers with codeFew-Shot Semantic Segmentation
13 benchmarks · 102 papers with code
Few-Shot Audio Classification
10 benchmarks · 5 papers with code
Cross-Domain Few-Shot
9 benchmarks · 80 papers with code
Few-Shot Relation Classification
4 benchmarks · 10 papers with code
One-Shot Learning
1 benchmark · 107 papers with code
5 shown of 12 sub-tasks (3 filed under another area). All sub-tasks of Few-Shot Learning →
Audio Classification
22 benchmarks · 183 papers with codeEnvironmental Sound Classification
3 benchmarks · 27 papers with code
Audio Multiple Target Classification
0 benchmarks · 1 paper with code
Parkinson Detection from Speech
0 benchmarks · 1 paper with code
Semi-supervised Audio Classification
0 benchmarks · 1 paper with code
4 shown of 4 sub-tasks.
Speech Enhancement
17 benchmarks · 280 papers with codeBandwidth Extension
6 benchmarks · 18 papers with code
Speech Dereverberation
5 benchmarks · 20 papers with code
Packet Loss Concealment
0 benchmarks · 4 papers with code
Speech Intelligibility Evaluation
0 benchmarks · 0 papers with code
4 shown of 4 sub-tasks.
2D Semantic Segmentation
11 benchmarks · 49 papers with codeImage Segmentation
13 benchmarks · 2,073 papers with code
Human Part Segmentation
6 benchmarks · 15 papers with code
Reflection Removal
5 benchmarks · 38 papers with code
Continual Semantic Segmentation
3 benchmarks · 17 papers with code
Text Style Transfer
2 benchmarks · 93 papers with code
5 shown of 17 sub-tasks (6 filed under another area). All sub-tasks of 2D Semantic Segmentation →
Emotion Recognition
9 benchmarks · 614 papers with codeSpeech Emotion Recognition
16 benchmarks · 139 papers with code
Emotion Recognition in Conversation
16 benchmarks · 83 papers with code
Multimodal Emotion Recognition
7 benchmarks · 80 papers with code
Emotion Recognition in Context
4 benchmarks · 5 papers with code
EEG Emotion Recognition
3 benchmarks · 14 papers with code
5 shown of 12 sub-tasks (2 filed under another area). All sub-tasks of Emotion Recognition →
DeepFake Detection
8 benchmarks · 237 papers with codeAudio Deepfake Detection
2 benchmarks · 35 papers with code
Multimodal Forgery Detection
1 benchmark · 1 paper with code
Synthetic Speech Detection
0 benchmarks · 12 papers with code
diffusion-generated faces detection
0 benchmarks · 1 paper with code
Human Detection of Deepfakes
0 benchmarks · 1 paper with code
5 shown of 5 sub-tasks.
Language Identification
6 benchmarks · 143 papers with codeNative Language Identification
1 benchmark · 5 papers with code
Dialect Identification
0 benchmarks · 33 papers with code
2 shown of 2 sub-tasks.
Text-To-Speech Synthesis
6 benchmarks · 104 papers with codeProsody Prediction
1 benchmark · 4 papers with code
Zero-Shot Multi-Speaker TTS
0 benchmarks · 3 papers with code
2 shown of 2 sub-tasks.
Speech Synthesis
5 benchmarks · 366 papers with codeSpeech Synthesis - Gujarati
2 benchmarks · 2 papers with code
Speech Synthesis - Assamese
1 benchmark · 1 paper with code
Speech Synthesis - Bengali
1 benchmark · 1 paper with code
Speech Synthesis - Bodo
1 benchmark · 1 paper with code
Speech Synthesis - Hindi
1 benchmark · 1 paper with code
5 shown of 15 sub-tasks. All sub-tasks of Speech Synthesis →
Text to Audio Retrieval
4 benchmarks · 13 papers with codeaudio moment retrieval
0 benchmarks · 2 papers with code
1 shown of 1 sub-task.
Accented Speech Recognition
4 benchmarks · 6 papers with codeSpeech Synthesis
5 benchmarks · 366 papers with code
1 shown of 1 sub-task.
Audio Generation
3 benchmarks · 124 papers with codeAudio Super-Resolution
4 benchmarks · 16 papers with code
Video-to-Sound Generation
1 benchmark · 6 papers with code
Voice Cloning
0 benchmarks · 36 papers with code
Room Impulse Response (RIR)
0 benchmarks · 18 papers with code
4 shown of 4 sub-tasks.
Environmental Sound Classification
3 benchmarks · 27 papers with codeSelf-Supervised Sound Classification
0 benchmarks · 1 paper with code
1 shown of 1 sub-task.
10-shot image generation
3 benchmarks · 21 papers with codeSemantic Segmentation
150 benchmarks · 6,644 papers with code
Text-to-Image Generation
17 benchmarks · 546 papers with code
Deblurring
17 benchmarks · 424 papers with code
Motion Synthesis
13 benchmarks · 126 papers with code
Image Deblurring
9 benchmarks · 167 papers with code
5 shown of 21 sub-tasks (14 filed under another area). All sub-tasks of 10-shot image generation →
Target Sound Extraction
3 benchmarks · 8 papers with codeStreaming Target Sound Extraction
1 benchmark · 1 paper with code
1 shown of 1 sub-task.
Online Beat Tracking
3 benchmarks · 4 papers with codeInference Optimization
0 benchmarks · 21 papers with code
1 shown of 1 sub-task.
Audio captioning
2 benchmarks · 62 papers with codeZero-shot Audio Captioning
2 benchmarks · 4 papers with code
Retrieval-augmented Few-shot In-context Audio Captioning
1 benchmark · 5 papers with code
2 shown of 2 sub-tasks.
Audio Source Separation
2 benchmarks · 54 papers with codeTarget Sound Extraction
3 benchmarks · 8 papers with code
Directional Hearing
1 benchmark · 1 paper with code
Single-Label Target Sound Extraction
0 benchmarks · 0 papers with code
3 shown of 3 sub-tasks.
Music Generation
1 benchmark · 190 papers with codeMusic Performance Rendering
0 benchmarks · 5 papers with code
Multimodal Music Generation
0 benchmarks · 3 papers with code
Music Texture Transfer
0 benchmarks · 1 paper with code
3 shown of 3 sub-tasks (1 filed under another area).
3D Human Pose Tracking
1 benchmark · 4 papers with codeMotion Synthesis
13 benchmarks · 126 papers with code
1 shown of 1 sub-task (1 filed under another area).
1 Image, 2*2 Stitchi
1 benchmark · 3 papers with codePose Estimation
31 benchmarks · 1,679 papers with code
Text-to-Image Generation
17 benchmarks · 546 papers with code
Image Deblurring
9 benchmarks · 167 papers with code
Virtual Try-on
9 benchmarks · 114 papers with code
Style Transfer
3 benchmarks · 759 papers with code
5 shown of 12 sub-tasks (8 filed under another area). All sub-tasks of 1 Image, 2*2 Stitchi →
Directional Hearing
1 benchmark · 1 paper with codeReal-time Directional Hearing
1 benchmark · 1 paper with code
1 shown of 1 sub-task.
Audio Signal Processing
0 benchmarks · 26 papers with codeblind source separation
0 benchmarks · 52 papers with code
Audio Compression
0 benchmarks · 17 papers with code
Audio Effects Modeling
0 benchmarks · 3 papers with code
3 shown of 3 sub-tasks.
Instance Search
0 benchmarks · 9 papers with codeAudio Fingerprint
0 benchmarks · 1 paper with code
1 shown of 1 sub-task.
3D Human Dynamics
0 benchmarks · 5 papers with codePortrait Animation
0 benchmarks · 13 papers with code
1 shown of 1 sub-task (1 filed under another area).
Audio Effects Modeling
0 benchmarks · 3 papers with codePitch control
0 benchmarks · 5 papers with code
Timbre Interpolation
0 benchmarks · 1 paper with code
2 shown of 2 sub-tasks.
Audio Signal Recognition
0 benchmarks · 2 papers with codeGunshot Detection
0 benchmarks · 1 paper with code
1 shown of 1 sub-task.
2D Classification
0 benchmarks · 0 papers with codeObject Detection
123 benchmarks · 4,657 papers with code
Deblurring
17 benchmarks · 424 papers with code
2D Pose Estimation
9 benchmarks · 46 papers with code
Anomaly Classification
5 benchmarks · 33 papers with code
Cell Detection
4 benchmarks · 60 papers with code
5 shown of 27 sub-tasks (8 filed under another area). All sub-tasks of 2D Classification →
Bird Classification
0 benchmarks · 0 papers with codeBird Audio Detection
0 benchmarks · 3 papers with code
Bird Species Classification With Audio-Visual Data
0 benchmarks · 0 papers with code
2 shown of 2 sub-tasks.
Hearing Aid and device processing
0 benchmarks · 0 papers with codeCadenza 1 - Task 1 - Headphone
1 benchmark · 1 paper with code
Cadenza 1 - Task 2 - In Car
1 benchmark · 1 paper with code
2 shown of 2 sub-tasks.
Signal Processing
0 benchmarks · 0 papers with codePhysiological Computing
0 benchmarks · 6 papers with code
1 shown of 1 sub-task (1 filed under another area).
Tasks with no parent task
36 tasks in Audio sit at the top of the archive's task tree with no sub-tasks of their own, most benchmarks first, then most papers with code.
Beat Tracking
15 benchmarks · 13 papers with code
Downbeat Tracking
13 benchmarks · 8 papers with code
Sound Event Detection
5 benchmarks · 92 papers with code
Acoustic Scene Classification
5 benchmarks · 42 papers with code
Sound Event Localization and Detection
5 benchmarks · 35 papers with code
Instrument Recognition
3 benchmarks · 26 papers with code
Voice Anti-spoofing
3 benchmarks · 15 papers with code
Audio Denoising
3 benchmarks · 10 papers with code
Text-to-Music Generation
2 benchmarks · 25 papers with code
Audio Tagging
1 benchmark · 49 papers with code
Sound Source Localization
1 benchmark · 39 papers with code
Direction of Arrival Estimation
1 benchmark · 17 papers with code
Lung Sound Classification
1 benchmark · 10 papers with code
Audio Quality Assessment
1 benchmark · 6 papers with code
fake voice detection
1 benchmark · 2 papers with code
Acoustic Novelty Detection
1 benchmark · 1 paper with code
Active Speaker Localization
1 benchmark · 0 papers with code
Sound Classification
0 benchmarks · 61 papers with code
X-ray Classification
0 benchmarks · 27 papers with code
Audio inpainting
0 benchmarks · 12 papers with code
Audio-Visual Synchronization
0 benchmarks · 11 papers with code
Chord Recognition
0 benchmarks · 7 papers with code
Visually Guided Sound Source Separation
0 benchmarks · 5 papers with code
Vowel Classification
0 benchmarks · 5 papers with code
Audio declipping
0 benchmarks · 4 papers with code
Music Genre Transfer
0 benchmarks · 4 papers with code
Music Compression
0 benchmarks · 3 papers with code
Music Quality Assessment
0 benchmarks · 2 papers with code
Audio Dequantization
0 benchmarks · 1 paper with code
Semi-Supervised Audio Regression
0 benchmarks · 1 paper with code
Shooter Localization
0 benchmarks · 1 paper with code
Soundscape evaluation
0 benchmarks · 1 paper with code
Speaker Orientation
0 benchmarks · 1 paper with code
Synthetic Song Detection
0 benchmarks · 1 paper with code
When should a hot water tank be replaced?
0 benchmarks · 1 paper with code
Video/Text-to-Audio Generation
0 benchmarks · 0 papers with code
2 tasks in Audio are filed only under a parent task from another area and are not listed on this page; the parent's task page carries them.
Task tree and counts are the archive's, frozen 2025-07-28 archive 2025-07-28. Nothing here is re-ranked.