{"url":"/dataset/crema-d","name":"CREMA-D","full_name":"CREMA-D","description_markdown":"**CREMA-D** is an emotional multimodal actor data set of 7,442 original clips from 91 actors. These clips were from 48 male and 43 female actors between the ages of 20 and 74 coming from a variety of races and ethnicities (African America, Asian, Caucasian, Hispanic, and Unspecified).\r\n\r\nActors spoke from a selection of 12 sentences. The sentences were presented using one of six different emotions (Anger, Disgust, Fear, Happy, Neutral, and Sad) and four different emotion levels (Low, Medium, High, and Unspecified).\r\n\r\nParticipants rated the emotion and emotion levels based on the combined audiovisual presentation, the video alone, and the audio alone. Due to the large number of ratings needed, this effort was crowd-sourced and a total of 2443 participants each rated 90 unique clips, 30 audio, 30 visual, and 30 audio-visual. 95% of the clips have more than 7 ratings.","description_withheld":null,"homepage":"https://github.com/CheyneyComputerScience/CREMA-D","introduced_date":null,"introduced_date_note":null,"introduced_by":null,"license":null,"modalities":[{"name":"Audio","url":"/datasets/modality/audio"}],"tasks":[{"name":"Facial Expression Recognition (FER)","url":"/task/facial-expression-recognition","datasets_with_task":"/datasets/task/facial-expression-recognition"},{"name":"Audio Classification","url":"/task/audio-classification","datasets_with_task":"/datasets/task/audio-classification"},{"name":"Speech Emotion Recognition","url":"/task/speech-emotion-recognition","datasets_with_task":"/datasets/task/speech-emotion-recognition"},{"name":"Self-Supervised Learning","url":"/task/self-supervised-learning","datasets_with_task":"/datasets/task/self-supervised-learning"},{"name":"Few-Shot Audio Classification","url":"/task/few-shot-audio-classification","datasets_with_task":"/datasets/task/few-shot-audio-classification"},{"name":"Talking Face Generation","url":"/task/talking-face-generation","datasets_with_task":"/datasets/task/talking-face-generation"},{"name":"Video Emotion Recognition","url":"/task/video-emotion-recognition","datasets_with_task":"/datasets/task/video-emotion-recognition"}],"languages":[],"variants":["CREMA-D"],"data_loaders":[{"repo":"https://github.com/tensorflow/datasets","url":"https://www.tensorflow.org/datasets/catalog/crema_d","frameworks":["tf","jax"]},{"repo":"https://github.com/CheyneyComputerScience/CREMA-D","url":"https://github.com/CheyneyComputerScience/CREMA-D","frameworks":[]}],"num_papers_in_archive":28,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/speech-emotion-recognition-on-crema-d","task":"Speech Emotion Recognition","dataset_variant":"CREMA-D","rows":9,"metrics":["Accuracy"],"first_row_in_archive_order":{"model":"Vertically long patch ViT","paper":"/paper/accuracy-enhancement-method-for-speech","metrics":{"Accuracy":"94.07"},"code_links":[{"title":"kjy7567/speech_emotion_recognition_from_log_Mel_spectrogram_using_vertically_long_patch","url":"https://github.com/kjy7567/speech_emotion_recognition_from_log_Mel_spectrogram_using_vertically_long_patch"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/video-emotion-recognition-on-crema-d","task":"Video Emotion Recognition","dataset_variant":"CREMA-D","rows":4,"metrics":["Accuracy","WAR"],"first_row_in_archive_order":{"model":"VAEmo","paper":"/paper/vaemo-efficient-representation-learning-for","metrics":{"Accuracy":"85.68%"},"code_links":[{"title":"MSA-LMC/VAEmo","url":"https://github.com/MSA-LMC/VAEmo"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/audio-classification-on-crema-d","task":"Audio Classification","dataset_variant":"CREMA-D","rows":3,"metrics":["Accuracy"],"first_row_in_archive_order":{"model":"EfficientLEAF","paper":"/paper/efficientleaf-a-faster-learnable-audio","metrics":{"Accuracy":"60.2"},"code_links":[{"title":"cpjku/efficientleaf","url":"https://github.com/cpjku/efficientleaf"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/few-shot-audio-classification-on-crema-d","task":"Few-Shot Audio Classification","dataset_variant":"CREMA-D","rows":3,"metrics":["Top-1 Accuracy(5-Way-1-Shot)"],"first_row_in_archive_order":{"model":"MT-SLVR (SimCLR + MLAP) w/ Parallel Adapters (FSD50K, RN18)","paper":"/paper/mt-slvr-multi-task-self-supervised-learning","metrics":{"Top-1 Accuracy(5-Way-1-Shot)":"29.61±0.38"},"code_links":[{"title":"cheggan/mt-slvr","url":"https://github.com/cheggan/mt-slvr"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/facial-expression-recognition-on-crema-d","task":"Facial Expression Recognition (FER)","dataset_variant":"CREMA-D","rows":1,"metrics":["UAR"],"first_row_in_archive_order":{"model":"EmoAffectNet LSTM","paper":"/paper/in-search-of-a-robust-facial-expressions","metrics":{"UAR":"79.0"},"code_links":[{"title":"ElenaRyumina/EMO-AffectNetModel","url":"https://github.com/ElenaRyumina/EMO-AffectNetModel"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/self-supervised-learning-on-crema-d","task":"Self-Supervised Learning","dataset_variant":"CREMA-D","rows":1,"metrics":["Accuracy"],"first_row_in_archive_order":{"model":"Hybrid BYOL-S/CvT","paper":"/paper/byol-s-learning-self-supervised-speech","metrics":{"Accuracy":"67.2"},"code_links":[{"title":"gasserelbanna/serab-byols","url":"https://github.com/gasserelbanna/serab-byols"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/talking-face-generation-on-crema-d","task":"Talking Face Generation","dataset_variant":"CREMA-D","rows":1,"metrics":["EmoAcc","FID","LSE-C"],"first_row_in_archive_order":{"model":"EmoGen","paper":"/paper/emotionally-enhanced-talking-face-generation","metrics":{"EmoAcc":"83.2","FID":"5.29","LSE-C":"6.663"},"code_links":[{"title":"sahilg06/EmoGen","url":"https://github.com/sahilg06/EmoGen"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/vaemo-efficient-representation-learning-for","title":"VAEmo: Efficient Representation Learning for Visual-Audio Emotion with Knowledge Injection","date":"2025-05-05","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/mtcae-dfer-multi-task-cascaded-autoencoder","title":"MTCAE-DFER: Multi-Task Cascaded Autoencoder for Dynamic Facial Expression Recognition","date":"2024-12-25","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/multimae-der-multimodal-masked-autoencoder","title":"MultiMAE-DER: Multimodal Masked Autoencoder for Dynamic Emotion Recognition","date":"2024-04-28","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/accuracy-enhancement-method-for-speech","title":"Accuracy enhancement method for speech emotion recognition from spectrogram using temporal frequency correlation and positional information learning through knowledge transfer","date":"2024-03-26","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/mt-slvr-multi-task-self-supervised-learning","title":"MT-SLVR: Multi-Task Self-Supervised Learning for Transformation In(Variant) Representations","date":"2023-05-29","rows_on_this_dataset":3,"code_links":1,"syntology":null},{"paper":"/paper/versatile-audio-visual-learning-for-handling","title":"Versatile audio-visual learning for emotion recognition","date":"2023-05-12","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/emotionally-enhanced-talking-face-generation","title":"Emotionally Enhanced Talking Face Generation","date":"2023-03-21","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/coordvit-a-novel-method-of-improve-vision","title":"CoordViT: A Novel Method of Improve Vision Transformer-Based Speech Emotion Recognition using Coordinate Information Concatenate","date":"2023-03-10","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/in-search-of-a-robust-facial-expressions","title":"In Search of a Robust Facial Expressions Recognition Model: A Large-Scale Visual Cross-Corpus Study","date":"2022-10-07","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/efficientleaf-a-faster-learnable-audio","title":"EfficientLEAF: A Faster LEarnable Audio Frontend of Questionable Use","date":"2022-07-12","rows_on_this_dataset":3,"code_links":1,"syntology":null},{"paper":"/paper/byol-s-learning-self-supervised-speech","title":"BYOL-S: Learning Self-supervised Speech Representations by Bootstrapping","date":"2022-06-24","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":7,"samples_ran":4,"samples_unverified":3,"pointer_only_for_licence":1,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/lerac-learning-rate-curriculum","title":"Learning Rate Curriculum","date":"2022-05-18","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/septr-separable-transformer-for-audio","title":"SepTr: Separable Transformer for Audio Spectrogram Processing","date":"2022-03-17","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/bigssl-exploring-the-frontier-of-large-scale","title":"BigSSL: Exploring the Frontier of Large-Scale Semi-Supervised Learning for Automatic Speech Recognition","date":"2021-09-27","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/ast-audio-spectrogram-transformer","title":"AST: Audio Spectrogram Transformer","date":"2021-04-05","rows_on_this_dataset":1,"code_links":5,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":3,"samples_ran":0,"samples_unverified":3,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/self-paced-ensemble-learning-for-speech-and","title":"Self-paced ensemble learning for speech and audio classification","date":"2021-03-22","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/non-linear-neurons-with-human-like-apical","title":"Non-linear Neurons with Human-like Apical Dendrite Activations","date":"2020-02-02","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/visually-guided-self-supervised-learning-of","title":"Visually Guided Self Supervised Learning of Speech Representations","date":"2020-01-13","rows_on_this_dataset":1,"code_links":0,"syntology":null}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":2,"samples_harvested":10,"samples_ran":4,"samples_unverified":6,"pointer_only_for_licence":1,"papers_with_no_sample_that_ran":1,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}