Papers › A Fine-tuned Wav2vec 2.0/HuBERT Benchmark For Speech Emotion Recognition, Speaker...

A Fine-tuned Wav2vec 2.0/HuBERT Benchmark For Speech Emotion Recognition, Speaker Verification and Spoken Language Understanding

4 Nov 2021arXiv:2111.02735archive 2025-07-28

Yingzhi Wang, Abdelmoumene Boumadane, Abdelwahab Heba

Speech self-supervised models such as wav2vec 2.0 and HuBERT are making revolutionary progress in Automatic Speech Recognition (ASR). However, they have not been totally proven to produce better performance on tasks other than ASR. In this work, we explored partial fine-tuning and entire fine-tuning on wav2vec 2.0 and HuBERT pre-trained models for three non-ASR speech tasks: Speech Emotion Recognition, Speaker Verification and Spoken Language Understanding. With simple proposed downstream frameworks, the best scores reached 79.58% weighted accuracy on speaker-dependent setting and 73.01% weighted accuracy on speaker-independent setting for Speech Emotion Recognition on IEMOCAP, 2.36% equal error rate for Speaker Verification on VoxCeleb1, 89.38% accuracy for Intent Classification and 78.92% F1 for Slot Filling on SLURP, showing the strength of fine-tuned wav2vec 2.0 and HuBERT on learning prosodic, voice-print and semantic representations.

PaperPDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Emotion RecognitionIntent ClassificationSlot FillingSpeaker VerificationSpeech Emotion RecognitionSpeech RecognitionSpoken Language Understandingintent-classificationslot-fillingspeech-recognition

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Intent Classification SLURP Partially Fine-tuned HuBERT Accuracy (%) 87.51 #2 of 5 Archive leaderboard report
Slot Filling SLURP Partially Fine-tuned HuBERT F1 0.753 #2 of 5 Archive leaderboard report
Speaker Verification VoxCeleb1 Fine-tuned HuBERT Large EER 2.36 #16 of 16 Archive leaderboard report
Speech Emotion Recognition IEMOCAP Partially Fine-tuned HuBERT Large WA 0.796 #4 of 8 Archive leaderboard report
Speech Emotion Recognition IEMOCAP Partially Fine-tuned HuBERT Large WA CV 0.730 #4 of 8 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections