{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/emotion-recognition-in-speech-using-cross","title":"Emotion Recognition in Speech using Cross-Modal Transfer in the Wild","arxiv_id":"1808.05561","date":"2018-08-16","proceeding":null,"authors":["Samuel Albanie","Arsha Nagrani","Andrea Vedaldi","Andrew Zisserman"],"abstract":"Obtaining large, human labelled speech datasets to train models for emotion\nrecognition is a notoriously challenging task, hindered by annotation cost and\nlabel ambiguity. In this work, we consider the task of learning embeddings for\nspeech classification without access to any form of labelled audio. We base our\napproach on a simple hypothesis: that the emotional content of speech\ncorrelates with the facial expression of the speaker. By exploiting this\nrelationship, we show that annotations of expression can be transferred from\nthe visual domain (faces) to the speech domain (voices) through cross-modal\ndistillation. We make the following contributions: (i) we develop a strong\nteacher network for facial emotion recognition that achieves the state of the\nart on a standard benchmark; (ii) we use the teacher to train a student, tabula\nrasa, to learn representations (embeddings) for speech emotion recognition\nwithout access to labelled audio data; and (iii) we show that the speech\nemotion embedding can be used for speech emotion recognition on external\nbenchmark datasets. Code, models and data are available.","url_abs":"http://arxiv.org/abs/1808.05561v1","url_pdf":"http://arxiv.org/pdf/1808.05561v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"emotion-recognition","task_name":"Emotion Recognition"},{"task_slug":"facial-emotion-recognition","task_name":"Facial Emotion Recognition"},{"task_slug":"facial-expression-recognition","task_name":"Facial Expression Recognition (FER)"},{"task_slug":"speech-emotion-recognition","task_name":"Speech Emotion Recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/facial-expression-recognition-on-ferplus","task":"Facial Expression Recognition (FER)","dataset":"FERPlus","model":"SENet Teacher","rank_in_archive_order":3,"of":4,"metrics":{"Accuracy(pretrained)":"88.88"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1808.05561","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}