Datasets › CREMA-D

CREMA-D

archive 2025-07-28

CREMA-D is an emotional multimodal actor data set of 7,442 original clips from 91 actors. These clips were from 48 male and 43 female actors between the ages of 20 and 74 coming from a variety of races and ethnicities (African America, Asian, Caucasian, Hispanic, and Unspecified).

Actors spoke from a selection of 12 sentences. The sentences were presented using one of six different emotions (Anger, Disgust, Fear, Happy, Neutral, and Sad) and four different emotion levels (Low, Medium, High, and Unspecified).

Participants rated the emotion and emotion levels based on the combined audiovisual presentation, the video alone, and the audio alone. Due to the large number of ratings needed, this effort was crowd-sourced and a total of 2443 participants each rated 90 unique clips, 30 audio, 30 visual, and 30 audio-visual. 95% of the clips have more than 7 ratings.

Benchmarks archive 2025-07-28

All 7 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

Papers archive 2025-07-28

18 shown of 18 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 28. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
VAEmo: Efficient Representation Learning for Visual-Audio Emotion with Knowledge Injection 1 1 5 May 2025 not harvested
MTCAE-DFER: Multi-Task Cascaded Autoencoder for Dynamic Facial Expression Recognition 1 1 25 Dec 2024 not harvested
MultiMAE-DER: Multimodal Masked Autoencoder for Dynamic Emotion Recognition 1 1 28 Apr 2024 not harvested
Accuracy enhancement method for speech emotion recognition from spectrogram using temporal frequency correlation and positional information learning through knowledge transfer 1 1 26 Mar 2024 not harvested
MT-SLVR: Multi-Task Self-Supervised Learning for Transformation In(Variant) Representations 1 3 29 May 2023 not harvested
Versatile audio-visual learning for emotion recognition 0 1 12 May 2023 not harvested
Emotionally Enhanced Talking Face Generation 1 1 21 Mar 2023 not harvested
CoordViT: A Novel Method of Improve Vision Transformer-Based Speech Emotion Recognition using Coordinate Information Concatenate 0 1 10 Mar 2023 not harvested
In Search of a Robust Facial Expressions Recognition Model: A Large-Scale Visual Cross-Corpus Study 1 1 7 Oct 2022 not harvested
EfficientLEAF: A Faster LEarnable Audio Frontend of Questionable Use 1 3 12 Jul 2022 not harvested
BYOL-S: Learning Self-supervised Speech Representations by Bootstrapping 1 1 24 Jun 2022 ran 4 of 7 samples (3 unverified; 1 pointer-only for licence)
Learning Rate Curriculum 1 1 18 May 2022 not harvested
SepTr: Separable Transformer for Audio Spectrogram Processing 1 1 17 Mar 2022 not harvested
BigSSL: Exploring the Frontier of Large-Scale Semi-Supervised Learning for Automatic Speech Recognition 0 1 27 Sep 2021 not harvested
AST: Audio Spectrogram Transformer 5 1 5 Apr 2021 ran 0 of 3 samples (3 unverified)
Self-paced ensemble learning for speech and audio classification 0 1 22 Mar 2021 not harvested
Non-linear Neurons with Human-like Apical Dendrite Activations 1 1 2 Feb 2020 not harvested
Visually Guided Self Supervised Learning of Speech Representations 0 1 13 Jan 2020 not harvested

Dataset loaders archive 2025-07-28

2 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • CREMA-D

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections