Home › Datasets › task › Speech Emotion Recognition
Speech Emotion Recognition datasets
archive 2025-07-28
22 datasets carry the task tag "Speech Emotion Recognition" (the task itself: Speech Emotion Recognition), ordered by the archive's paper count. Page 1 of 1: 22 shown of 22. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Speech Emotion Recognition datasets 1–22 of 22
The MNIST database (Modified National Institute of Standards and Technology database) is a large collection of handwritten digits.
7,651 papers · 44 benchmarks
IEMOCAP (The Interactive Emotional Dyadic Motion Capture (IEMOCAP) Database)
Multimodal Emotion Recognition IEMOCAP The IEMOCAP dataset consists of 151 videos of recorded dialogues, with 2 speakers per session for a total of 302 videos across the dataset.
749 papers · 3 benchmarks
MELD (Multimodal EmotionLines Dataset)
Multimodal EmotionLines Dataset (MELD) has been created by enhancing and extending EmotionLines dataset.
289 papers · 3 benchmarks
MSP-IMPROV (MSP-IMPROV: An Acted Corpus of Dyadic Interactions to Study Emotion Perception)
We present the MSP-IMPROV corpus, a multimodal emotional database, where the goal is to have control over lexical content and emotion while also promoting naturalness in the recordings.
70 papers · 1 benchmark
CREMA-D is an emotional multimodal actor data set of 7,442 original clips from 91 actors.
28 papers · 7 benchmarks
RAVDESS (Ryerson Audio-Visual Database of Emotional Speech and Song)
The Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS) contains 7,356 files (total size: 24.8 GB).
27 papers · 6 benchmarks
ShEMO (Sharif Emotional Speech Database)
The database includes 3000 semi-natural utterances, equivalent to 3 hours and 25 minutes of speech data extracted from online radio plays.
16 papers · 1 benchmark
MSP-Podcast (A large naturalistic speech emotional dataset)
The MSP-Podcast corpus contains speech segments from podcast recordings which are perceptually annotated using crowdsourcing.
9 papers · 4 benchmarks
The EMODB database is the freely available German emotional database.
6 papers · 1 benchmark
LSSED, a challenging large-scale english dataset for speech emotion recognition.
6 papers · 1 benchmark
RESD (Russian Emotional Speech Dialogs with annotated text)
Russian dataset of emotional speech dialogues.
3 papers · 1 benchmark
A database of more than 2000 minutes of audio-visual data of 398 people coming from six cultures, 50% female, and uniformly spanning the age range of 18 to 65 years old.
3 papers · 0 benchmarks
Dusha (Dusha Crowd, Dusha Podcast)
Dusha is a dataset for speech emotion recognition (SER) tasks.
2 papers · 2 benchmarks
Overview nEMO is a simulated dataset of emotional speech in the Polish language.
2 papers · 0 benchmarks
AESI (Athens Emotional States Inventory)
The development of ecologically valid procedures for collecting reliable and unbiased emotional data towards computer interfaces with social and affective intelligence targeting patients with mental disorders.
1 paper · 0 benchmarks
BERSt (Basic Emotion Random phrase Shouts)
BERSt Dataset We release the BERSt Dataset for various speech recognition tasks including Automatic Speech Recognition (ASR) and Speech Emotion Recogniton (SER) Overview 4526 single phrase recordings (~3.75h) 98 professional actors 19…
1 paper · 1 benchmark
EmoFilm (Emotional speech from Films)
EmoFilm is a multilingual emotional speech corpus comprising 1115 audio instances produced in English, Italian, and Spanish languages.
1 paper · 0 benchmarks
Quechua Collao corpus for automatic emotion recognition in speech.
1 paper · 1 benchmark
SES (Spanish Emotional Speech)
Currently, an essential point in speech synthesis is the addressing of the variability of human speech.
1 paper · 0 benchmarks
VNEMOS (Vietnamese Speech Emotion Dataset)
This research introduces the dataset that we created to test voice emotional recognition models with Vietnamese data.
1 paper · 0 benchmarks
A modification on the ShEMO dataset with help of an Automatic Speech Recognition (ASR) system.
1 paper · 0 benchmarks
JTES (Japanese Twitter-based Emotional Speech)
We designed an emotional speech database that can be used for emotion recognition as well as recognition and synthesis of speech with various emotions.
0 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.