{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/multimodal-speech-emotion-recognition-using","title":"Multimodal Speech Emotion Recognition Using Audio and Text","arxiv_id":"1810.04635","date":"2018-10-10","proceeding":null,"authors":["Seunghyun Yoon","Seokhyun Byun","Kyomin Jung"],"abstract":"Speech emotion recognition is a challenging task, and extensive reliance has\nbeen placed on models that use audio features in building well-performing\nclassifiers. In this paper, we propose a novel deep dual recurrent encoder\nmodel that utilizes text data and audio signals simultaneously to obtain a\nbetter understanding of speech data. As emotional dialogue is composed of sound\nand spoken content, our model encodes the information from audio and text\nsequences using dual recurrent neural networks (RNNs) and then combines the\ninformation from these sources to predict the emotion class. This architecture\nanalyzes speech data from the signal level to the language level, and it thus\nutilizes the information within the data more comprehensively than models that\nfocus on audio features. Extensive experiments are conducted to investigate the\nefficacy and properties of the proposed model. Our proposed model outperforms\nprevious state-of-the-art methods in assigning data to one of four emotion\ncategories (i.e., angry, happy, sad and neutral) when the model is applied to\nthe IEMOCAP dataset, as reflected by accuracies ranging from 68.8% to 71.8%.","url_abs":"http://arxiv.org/abs/1810.04635v1","url_pdf":"http://arxiv.org/pdf/1810.04635v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"multimodal-speech-emotion-recognition-using","repo_url":"https://github.com/MagnusXu/Speech-Emotion-Recognition-Capstone-Project","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null},{"paper_slug":"multimodal-speech-emotion-recognition-using","repo_url":"https://github.com/SER-2020-Project-ZX/Reference","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null},{"paper_slug":"multimodal-speech-emotion-recognition-using","repo_url":"https://github.com/aris-ai/Audio-and-text-based-emotion-recognition","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"multimodal-speech-emotion-recognition-using","repo_url":"https://github.com/david-yoon/multimodal-speech-emotion","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"emotion-classification","task_name":"Emotion Classification"},{"task_slug":"emotion-recognition","task_name":"Emotion Recognition"},{"task_slug":"multimodal-emotion-recognition","task_name":"Multimodal Emotion Recognition"},{"task_slug":"multimodal-sentiment-analysis","task_name":"Multimodal Sentiment Analysis"},{"task_slug":"speech-emotion-recognition","task_name":"Speech Emotion Recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1810.04635","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}