{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/wav2small-distilling-wav2vec2-to-72k-1","title":"Wav2Small: Distilling Wav2Vec2 to 72K parameters for Low-Resource Speech emotion recognition","arxiv_id":"2408.13920","date":"2024-08-25","proceeding":null,"authors":["Dionyssos Kounadis-Bastian","Oliver Schrüfer","Anna Derington","Hagen Wierstorf","Florian Eyben","Felix Burkhardt","Björn Schuller"],"abstract":"Speech Emotion Recognition (SER) needs high computational resources to overcome the challenge of substantial annotator disagreement. Today SER is shifting towards dimensional annotations of arousal, dominance, and valence (A/D/V). Universal metrics as the L2 distance prove unsuitable for evaluating A/D/V accuracy due to non converging consensus of annotator opinions. However, Concordance Correlation Coefficient (CCC) arose as an alternative metric for A/D/V where a model's output is evaluated to match a whole dataset's CCC rather than L2 distances of individual audios. Recent studies have shown that wav2vec2 / wavLM architectures outputing a float value for each A/D/V dimension achieve today's State-of-the-art (Sota) CCC on A/D/V. The Wav2Vec2.0 / WavLM family has a high computational footprint, but training small models using human annotations has been unsuccessful. In this paper we use a large Transformer Sota A/D/V model as Teacher/Annotator to train 5 student models: 4 MobileNets and our proposed Wav2Small, using only the Teacher's A/D/V outputs instead of human annotations. The Teacher model we propose also sets a new Sota on the MSP Podcast dataset of valence CCC=0.676. We choose MobileNetV4 / MobileNet-V3 as students, as MobileNet has been designed for fast execution times. We also propose Wav2Small - an architecture designed for minimal parameters and RAM consumption. Wav2Small with an .onnx (quantised) of only 120KB is a potential solution for A/D/V on hardware with low resources, having only 72K parameters vs 3.12M parameters for MobileNet-V4-Small.","url_abs":"https://arxiv.org/abs/2408.13920v4","url_pdf":"https://arxiv.org/pdf/2408.13920v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"links_only","authors_date_abstract":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license), from the Kaggle arXiv metadata snapshot of 2026-09-12"},"code_links":[{"paper_slug":"wav2small-distilling-wav2vec2-to-72k-1","repo_url":"https://github.com/dkounadis/wav2small","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/speech-emotion-recognition-on-msp-podcast-1","task":"Speech Emotion Recognition","dataset":"MSP-Podcast (Activation)","model":"wav2small-Teacher","rank_in_archive_order":1,"of":4,"metrics":{"CCC":"0.7620181"},"uses_additional_data":false},{"leaderboard":"/sota/speech-emotion-recognition-on-msp-podcast-2","task":"Speech Emotion Recognition","dataset":"MSP-Podcast (Dominance)","model":"wav2small-Teacher","rank_in_archive_order":1,"of":4,"metrics":{"CCC":"0.6840044"},"uses_additional_data":false},{"leaderboard":"/sota/speech-emotion-recognition-on-msp-podcast","task":"Speech Emotion Recognition","dataset":"MSP-Podcast (Valence)","model":"wav2small-Teacher","rank_in_archive_order":1,"of":4,"metrics":{"CCC":"0.676"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}