Papers › MultiMAE-DER: Multimodal Masked Autoencoder for Dynamic Emotion Recognition

MultiMAE-DER: Multimodal Masked Autoencoder for Dynamic Emotion Recognition

28 Apr 2024arXiv:2404.18327archive 2025-07-28

Peihao Xiang, Chaohao Lin, Kaida Wu, Ou Bai

This paper presents a novel approach to processing multimodal data for dynamic emotion recognition, named as the Multimodal Masked Autoencoder for Dynamic Emotion Recognition (MultiMAE-DER). The MultiMAE-DER leverages the closely correlated representation information within spatiotemporal sequences across visual and audio modalities. By utilizing a pre-trained masked autoencoder model, the MultiMAEDER is accomplished through simple, straightforward finetuning. The performance of the MultiMAE-DER is enhanced by optimizing six fusion strategies for multimodal input sequences. These strategies address dynamic feature correlations within cross-domain data across spatial, temporal, and spatiotemporal sequences. In comparison to state-of-the-art multimodal supervised learning models for dynamic emotion recognition, MultiMAE-DER enhances the weighted average recall (WAR) by 4.41% on the RAVDESS dataset and by 2.06% on the CREMAD. Furthermore, when compared with the state-of-the-art model of multimodal self-supervised learning, MultiMAE-DER achieves a 1.86% higher WAR on the IEMOCAP dataset.

PaperPDFCode

Code

Peihao-Xiang/MultiMAE-DFER officialmentioned on GitHubtf report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Emotion RecognitionMultimodal Emotion RecognitionSelf-Supervised LearningVideo Emotion Recognition

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Emotion Recognition RAVDESS MultiMAE-DER WAR 83.61% #5 of 5 Archive leaderboard report
Multimodal Emotion Recognition IEMOCAP-4 MultiMAE-DER Weighted Recall 63.73 #11 of 11 Archive leaderboard report
Video Emotion Recognition CREMA-D MultiMAE-DER WAR 79.36% #4 of 4 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

AttentionDenoising AutoencoderDense ConnectionsLayer NormalizationLinear LayerMulti-Head AttentionResidual ConnectionSelf-LearningSoftmaxVision Transformer

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections