{"url":"/dataset/msp-podcast","name":"MSP-Podcast","full_name":"A large naturalistic speech emotional dataset","description_markdown":"The MSP-Podcast corpus contains speech segments from podcast recordings which are perceptually annotated using crowdsourcing. The collection of this corpus is an ongoing process. Version 1.7 of the corpus has 62,140 speaking turns (100hrs).\r\n\r\nKey features of this corpus:\r\n\r\n* We download available audio recordings with common license. We only use the podcasts that have less restrictive licenses, so we can modify, sell and distribute the corpus (you can use it for commercial product!). \r\n* Most of the segments in a regular podcasts are neutral. We use machine learning techniques trained with available data to retrieve candidate segments. These segments are emotionally annotated with crowdsourcing. This approach allows us to spend our resources on speech segments that are likely to convey emotions.\r\n* We annotate categorical emotions and attribute based labels at the speaking turn label\r\n* This is an ongoing effort, where we currently have 62,140 speaking turns (100h). We collect approximately 10,000-13,000 new speaking turns per year. Our goal is to reach 400 hours.","description_withheld":null,"homepage":"https://ecs.utdallas.edu/research/researchlabs/msp-lab/MSP-Podcast.html","introduced_date":null,"introduced_date_note":null,"introduced_by":null,"license":null,"modalities":[{"name":"Audio","url":"/datasets/modality/audio"}],"tasks":[{"name":"Emotion Recognition","url":"/task/emotion-recognition","datasets_with_task":"/datasets/task/emotion-recognition"},{"name":"Speech Emotion Recognition","url":"/task/speech-emotion-recognition","datasets_with_task":"/datasets/task/speech-emotion-recognition"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["MSP-Podcast (Valence)","MSP-Podcast (Activation)","MSP-Podcast (Dominance)","MSP-Podcast"],"data_loaders":[],"num_papers_in_archive":9,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/speech-emotion-recognition-on-msp-podcast","task":"Speech Emotion Recognition","dataset_variant":"MSP-Podcast (Valence)","rows":4,"metrics":["CCC"],"first_row_in_archive_order":{"model":"wav2small-Teacher","paper":"/paper/wav2small-distilling-wav2vec2-to-72k-1","metrics":{"CCC":"0.676"},"code_links":[{"title":"dkounadis/wav2small","url":"https://github.com/dkounadis/wav2small"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/speech-emotion-recognition-on-msp-podcast-1","task":"Speech Emotion Recognition","dataset_variant":"MSP-Podcast (Activation)","rows":4,"metrics":["CCC"],"first_row_in_archive_order":{"model":"wav2small-Teacher","paper":"/paper/wav2small-distilling-wav2vec2-to-72k-1","metrics":{"CCC":"0.7620181"},"code_links":[{"title":"dkounadis/wav2small","url":"https://github.com/dkounadis/wav2small"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/speech-emotion-recognition-on-msp-podcast-2","task":"Speech Emotion Recognition","dataset_variant":"MSP-Podcast (Dominance)","rows":4,"metrics":["CCC"],"first_row_in_archive_order":{"model":"wav2small-Teacher","paper":"/paper/wav2small-distilling-wav2vec2-to-72k-1","metrics":{"CCC":"0.6840044"},"code_links":[{"title":"dkounadis/wav2small","url":"https://github.com/dkounadis/wav2small"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/emotion-recognition-on-msp-podcast","task":"Emotion Recognition","dataset_variant":"MSP-Podcast","rows":1,"metrics":["Concordance correlation coefficient (CCC)"],"first_row_in_archive_order":{"model":"w2v2-L-robust-12","paper":"/paper/dawn-of-the-transformer-era-in-speech-emotion","metrics":{"Concordance correlation coefficient (CCC)":"0.638"},"code_links":[{"title":"audeering/w2v2-how-to","url":"https://github.com/audeering/w2v2-how-to"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/odyssey-2024-speech-emotion-recognition","title":"Odyssey 2024 - Speech Emotion Recognition Challenge: Dataset, Baseline Framework, and Results","date":"2024-06-20","rows_on_this_dataset":3,"code_links":1,"syntology":null},{"paper":"/paper/dawn-of-the-transformer-era-in-speech-emotion","title":"Dawn of the transformer era in speech emotion recognition: closing the valence gap","date":"2022-03-14","rows_on_this_dataset":4,"code_links":1,"syntology":null},{"paper":"/paper/contrastive-unsupervised-learning-for-speech","title":"Contrastive Unsupervised Learning for Speech Emotion Recognition","date":"2021-02-12","rows_on_this_dataset":3,"code_links":0,"syntology":null},{"paper":"/paper/wav2small-distilling-wav2vec2-to-72k-1","title":"Wav2Small: Distilling Wav2Vec2 to 72K parameters for Low-Resource Speech emotion recognition","date":null,"rows_on_this_dataset":3,"code_links":1,"syntology":null}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}