Papers › A vector quantized masked autoencoder for speech emotion recognition

A vector quantized masked autoencoder for speech emotion recognition

21 Apr 2023arXiv:2304.11117archive 2025-07-28

Samir Sadok, Simon Leglaive, Renaud Séguier

Recent years have seen remarkable progress in speech emotion recognition (SER), thanks to advances in deep learning techniques. However, the limited availability of labeled data remains a significant challenge in the field. Self-supervised learning has recently emerged as a promising solution to address this challenge. In this paper, we propose the vector quantized masked autoencoder for speech (VQ-MAE-S), a self-supervised model that is fine-tuned to recognize emotions from speech signals. The VQ-MAE-S model is based on a masked autoencoder (MAE) that operates in the discrete latent space of a vector-quantized variational autoencoder. Experimental results show that the proposed VQ-MAE-S model, pre-trained on the VoxCeleb2 dataset and fine-tuned on emotional speech data, outperforms an MAE working on the raw spectrogram representation and other state-of-the-art methods in SER.

PaperPDFCode

Code

samsad35/VQ-MAE-S-code officialmentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Emotion RecognitionSelf-Supervised LearningSpeech Emotion Recognition

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Speech Emotion Recognition EmoDB Dataset VQ-MAE-S-12 (Frame) + Query2Emo Accuracy 90.2 #1 of 1 Archive leaderboard report
Speech Emotion Recognition EmoDB Dataset VQ-MAE-S-12 (Frame) + Query2Emo F1 0.891 #1 of 1 Archive leaderboard report
Speech Emotion Recognition RAVDESS VQ-MAE-S-12 (Frame) + Query2Emo Accuracy 84.1 #1 of 5 Archive leaderboard report
Speech Emotion Recognition RAVDESS VQ-MAE-S-12 (Frame) + Query2Emo F1 0.844 #1 of 5 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

MAE

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections