Datasets › CMU-MOSEI

CMU-MOSEI

Introduced by AmirAli Bagher Zadeh et al. in Multimodal Language Analysis in the Wild: CMU-MOSEI Dataset and Interpretable Dynamic Fusion Graph1 Jan 2018 archive 2025-07-28

CMU Multimodal Opinion Sentiment and Emotion Intensity (CMU-MOSEI) is the largest dataset of sentence-level sentiment analysis and emotion recognition in online videos. CMU-MOSEI contains over 12 hours of annotated video from over 1000 speakers and 250 topics.

Benchmarks archive 2025-07-28

All 3 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Multimodal Sentiment Analysis CMU-MOSEI SeMUL-PCD Accuracy 88.62 Multi-label Emotion Analysis in Conversation via... — 15 Compare
Emotion Classification CMU-MOSEI MARLIN (ViT-L) Accuracy 80.63 MARLIN: Masked Autoencoder for facial video... ControlNet/MARLIN 4 Compare
Facial Expression Recognition CMU-MOSEI ConCluGen Weighted Accuracy 66.48 Multi-Task Multi-Modal Self-Supervised Learning for... tub-cv-group/conclugen 1 Compare

Papers archive 2025-07-28

14 shown of 14 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 190. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Multi-Task Multi-Modal Self-Supervised Learning for Facial Expression Recognition 1 2 16 Apr 2024 not harvested
Multi-label Emotion Analysis in Conversation via Multimodal Knowledge Distillation 0 1 27 Oct 2023 not harvested
Learning Language-guided Adaptive Hyper-modality Representation for Multimodal Sentiment Analysis 1 1 9 Oct 2023 not harvested
Multimodal Multi-loss Fusion Network for Sentiment Analysis 1 1 1 Aug 2023 not harvested
Speech-Text Dialog Pre-training for Spoken Dialog Understanding with Explicit Cross-Modal Alignment 1 1 19 May 2023 not harvested
UniMSE: Towards Unified Multimodal Sentiment Analysis and Emotion Recognition 1 1 21 Nov 2022 not harvested
MARLIN: Masked Autoencoder for facial video Representation LearnINg 1 6 12 Nov 2022 not harvested
MMLatch: Bottom-up Top-down Fusion for Multimodal Sentiment Analysis 1 1 24 Jan 2022 not harvested
Unsupervised Multimodal Language Representations using Convolutional Autoencoders 0 1 6 Oct 2021 not harvested
Modulated Fusion using Transformer for Linguistic-Acoustic Emotion Recognition 1 1 5 Oct 2020 not harvested
A Transformer-based joint-encoding for Emotion Recognition and Sentiment Analysis 1 1 29 Jun 2020 ran 1 of 4 samples (3 unverified)
Gated Mechanism for Attention Based Multimodal Sentiment Analysis 0 1 21 Feb 2020 not harvested
Multilogue-Net: A Context Aware RNN for Multi-modal Emotion Detection and Sentiment Analysis in Conversation 1 1 19 Feb 2020 ran 2 of 4 samples (2 unverified)
Multimodal Language Analysis in the Wild: CMU-MOSEI Dataset and Interpretable Dynamic Fusion Graph 0 1 1 Jul 2018 not harvested

Dataset loaders archive 2025-07-28

1 loader as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

Custom

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • CMU-MOSEI

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections