Datasets › MSP-Podcast

MSP-Podcast (A large naturalistic speech emotional dataset)

archive 2025-07-28

The MSP-Podcast corpus contains speech segments from podcast recordings which are perceptually annotated using crowdsourcing. The collection of this corpus is an ongoing process. Version 1.7 of the corpus has 62,140 speaking turns (100hrs).

Key features of this corpus:

  • We download available audio recordings with common license. We only use the podcasts that have less restrictive licenses, so we can modify, sell and distribute the corpus (you can use it for commercial product!).
  • Most of the segments in a regular podcasts are neutral. We use machine learning techniques trained with available data to retrieve candidate segments. These segments are emotionally annotated with crowdsourcing. This approach allows us to spend our resources on speech segments that are likely to convey emotions.
  • We annotate categorical emotions and attribute based labels at the speaking turn label
  • This is an ongoing effort, where we currently have 62,140 speaking turns (100h). We collect approximately 10,000-13,000 new speaking turns per year. Our goal is to reach 400 hours.

Benchmarks archive 2025-07-28

All 4 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Speech Emotion Recognition MSP-Podcast (Valence) wav2small-Teacher CCC 0.676 Wav2Small: Distilling Wav2Vec2 to 72K parameters for... dkounadis/wav2small 4 Compare
Speech Emotion Recognition MSP-Podcast (Activation) wav2small-Teacher CCC 0.7620181 Wav2Small: Distilling Wav2Vec2 to 72K parameters for... dkounadis/wav2small 4 Compare
Speech Emotion Recognition MSP-Podcast (Dominance) wav2small-Teacher CCC 0.6840044 Wav2Small: Distilling Wav2Vec2 to 72K parameters for... dkounadis/wav2small 4 Compare
Emotion Recognition MSP-Podcast w2v2-L-robust-12 Concordance correlation coefficient (CCC) 0.638 Dawn of the transformer era in speech emotion... audeering/w2v2-how-to 1 Compare

Papers archive 2025-07-28

4 shown of 4 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 9. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • MSP-Podcast (Valence)
  • MSP-Podcast (Activation)
  • MSP-Podcast (Dominance)
  • MSP-Podcast

4 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections