{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-fine-tuned-wav2vec-2-0-hubert-benchmark-for","title":"A Fine-tuned Wav2vec 2.0/HuBERT Benchmark For Speech Emotion Recognition, Speaker Verification and Spoken Language Understanding","arxiv_id":"2111.02735","date":"2021-11-04","proceeding":null,"authors":["Yingzhi Wang","Abdelmoumene Boumadane","Abdelwahab Heba"],"abstract":"Speech self-supervised models such as wav2vec 2.0 and HuBERT are making revolutionary progress in Automatic Speech Recognition (ASR). However, they have not been totally proven to produce better performance on tasks other than ASR. In this work, we explored partial fine-tuning and entire fine-tuning on wav2vec 2.0 and HuBERT pre-trained models for three non-ASR speech tasks: Speech Emotion Recognition, Speaker Verification and Spoken Language Understanding. With simple proposed downstream frameworks, the best scores reached 79.58% weighted accuracy on speaker-dependent setting and 73.01% weighted accuracy on speaker-independent setting for Speech Emotion Recognition on IEMOCAP, 2.36% equal error rate for Speaker Verification on VoxCeleb1, 89.38% accuracy for Intent Classification and 78.92% F1 for Slot Filling on SLURP, showing the strength of fine-tuned wav2vec 2.0 and HuBERT on learning prosodic, voice-print and semantic representations.","url_abs":"https://arxiv.org/abs/2111.02735v3","url_pdf":"https://arxiv.org/pdf/2111.02735v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"automatic-speech-recognition-2","task_name":"Automatic Speech Recognition"},{"task_slug":"automatic-speech-recognition","task_name":"Automatic Speech Recognition (ASR)"},{"task_slug":"emotion-recognition","task_name":"Emotion Recognition"},{"task_slug":"intent-classification","task_name":"Intent Classification"},{"task_slug":"slot-filling","task_name":"Slot Filling"},{"task_slug":"speaker-verification","task_name":"Speaker Verification"},{"task_slug":"speech-emotion-recognition","task_name":"Speech Emotion Recognition"},{"task_slug":"speech-recognition","task_name":"Speech Recognition"},{"task_slug":"spoken-language-understanding","task_name":"Spoken Language Understanding"},{"task_slug":"intent-classification-1","task_name":"intent-classification"},{"task_slug":"slot-filling-1","task_name":"slot-filling"},{"task_slug":"speech-recognition-1","task_name":"speech-recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/intent-classification-on-slurp","task":"Intent Classification","dataset":"SLURP","model":"Partially Fine-tuned HuBERT","rank_in_archive_order":2,"of":5,"metrics":{"Accuracy (%)":"87.51"},"uses_additional_data":false},{"leaderboard":"/sota/slot-filling-on-slurp","task":"Slot Filling","dataset":"SLURP","model":"Partially Fine-tuned HuBERT","rank_in_archive_order":2,"of":5,"metrics":{"F1":"0.753"},"uses_additional_data":false},{"leaderboard":"/sota/speaker-verification-on-voxceleb1","task":"Speaker Verification","dataset":"VoxCeleb1","model":"Fine-tuned HuBERT Large","rank_in_archive_order":16,"of":16,"metrics":{"EER":"2.36"},"uses_additional_data":false},{"leaderboard":"/sota/speech-emotion-recognition-on-iemocap","task":"Speech Emotion Recognition","dataset":"IEMOCAP","model":"Partially Fine-tuned HuBERT Large","rank_in_archive_order":4,"of":8,"metrics":{"WA":"0.796","WA CV":"0.730"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=2111.02735","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}