Papers › Multi-Task Multi-Modal Self-Supervised Learning for Facial Expression Recognition

Multi-Task Multi-Modal Self-Supervised Learning for Facial Expression Recognition

16 Apr 2024arXiv:2404.10904archive 2025-07-28

Marah Halawa, Florian Blume, Pia Bideau, Martin Maier, Rasha Abdel Rahman, Olaf Hellwich

Human communication is multi-modal; e.g., face-to-face interaction involves auditory signals (speech) and visual signals (face movements and hand gestures). Hence, it is essential to exploit multiple modalities when designing machine learning-based facial expression recognition systems. In addition, given the ever-growing quantities of video data that capture human facial expressions, such systems should utilize raw unlabeled videos without requiring expensive annotations. Therefore, in this work, we employ a multitask multi-modal self-supervised learning method for facial expression recognition from in-the-wild video data. Our model combines three self-supervised objective functions: First, a multi-modal contrastive loss, that pulls diverse data modalities of the same video together in the representation space. Second, a multi-modal clustering loss that preserves the semantic structure of input data in the representation space. Finally, a multi-modal data reconstruction loss. We conduct a comprehensive study on this multimodal multi-task self-supervised learning method on three facial expression recognition benchmarks. To that end, we examine the performance of learning through different combinations of self-supervised tasks on the facial expression recognition downstream task. Our model ConCluGen outperforms several multi-modal self-supervised and fully supervised baselines on the CMU-MOSEI dataset. Our results generally show that multi-modal self-supervision tasks offer large performance gains for challenging tasks such as facial expression recognition, while also reducing the amount of manual annotations required. We release our pre-trained models as well as source code publicly

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

tub-cv-group/conclugen officialmentioned in papermentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Emotion ClassificationEmotion Recognition in ConversationFacial Expression RecognitionSelf-Supervised Learning

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Emotion Classification CMU-MOSEI ConCluGen Accuracy 66.48 #4 of 4 Archive leaderboard report
Emotion Classification CMU-MOSEI ConCluGen Weighted Accuracy 66.48 #4 of 4 Archive leaderboard report
Emotion Recognition in Conversation MELD ConCluGen Accuracy 60.03 #67 of 68 Archive leaderboard report
Emotion Recognition in Conversation MELD ConCluGen Balanced Accuracy 60.03 #67 of 68 Archive leaderboard report
Facial Expression Recognition CMU-MOSEI ConCluGen Weighted Accuracy 66.48 #1 of 1 Archive leaderboard report
Facial Expression Recognition MELD ConCluGen Weighted Accuracy 60.03 #1 of 1 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Contrastive Learningk-Means Clustering

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections