Datasets › RAVDESS
RAVDESS (Ryerson Audio-Visual Database of Emotional Speech and Song)
The Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS) contains 7,356 files (total size: 24.8 GB). The database contains 24 professional actors (12 female, 12 male), vocalizing two lexically-matched statements in a neutral North American accent. Speech includes calm, happy, sad, angry, fearful, surprise, and disgust expressions, and song contains calm, happy, sad, angry, and fearful emotions. Each expression is produced at two levels of emotional intensity (normal, strong), with an additional neutral expression. All conditions are available in three modality formats: Audio-only (16bit, 48kHz .wav), Audio-Video (720p H.264, AAC 48kHz, .mp4), and Video-only (no sound). Note, there are no song files for Actor_18.
Paper: The Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS): A dynamic, multimodal set of facial and vocal expressions in North American English Source:
Benchmarks archive 2025-07-28
All 6 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.
| First row (archive order) | Paper | Code | ||||
|---|---|---|---|---|---|---|
| Emotion Recognition | RAVDESS | LogisticRegression on posteriors of xlsr-Wav2Vec2.0&bi-LSTM+Attention Accuracy 86.70% | A proposal for Multimodal Emotion Recognition using... | cristinalunaj/MMEmotionRecognition | 5 | Compare |
| Speech Emotion Recognition | RAVDESS | VQ-MAE-S-12 (Frame) + Query2Emo Accuracy 84.1 | A vector quantized masked autoencoder for speech emotion... | samsad35/VQ-MAE-S-code | 5 | Compare |
| Facial Emotion Recognition | RAVDESS | MTCAE-DFER Accuracy 83.69% | MTCAE-DFER: Multi-Task Cascaded Autoencoder for Dynamic... | Peihao-Xiang/MTCAE-DFER | 4 | Compare |
| Audio Classification | RAVDESS | ASM-RH-A Top-1 Accuracy 75.4 | Mixer is more than just a model | — | 2 | Compare |
| Emotion Classification | RAVDESS | ERANN-0-4 Top-1 Accuracy 74.8 | ERANNs: Efficient Residual Audio Neural Networks for... | — | 1 | Compare |
| Facial Expression Recognition (FER) | RAVDESS | EmoAffectNet LSTM UAR 69.7 | In Search of a Robust Facial Expressions Recognition... | ElenaRyumina/EMO-AffectNetModel | 1 | Compare |
Papers archive 2025-07-28
9 shown of 9 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 27. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.
Dataset loaders archive 2025-07-28
1 loader as listed in the archive; links are outbound and not re-checked here.
Tasks archive 2025-07-28
License archive 2025-07-28
Attribution-NonCommercial-ShareAlike 4.0 International
Modalities archive 2025-07-28
Languages archive 2025-07-28
Variants archive 2025-07-28
- RAVDESS
1 variant name, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections