{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/190410788","title":"Speech Emotion Recognition Using Multi-hop Attention Mechanism","arxiv_id":"1904.10788","date":"2019-04-23","proceeding":"2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) 2019 5","authors":["Seunghyun Yoon","Seokhyun Byun","Subhadeep Dey","Kyomin Jung"],"abstract":"In this paper, we are interested in exploiting textual and acoustic data of\nan utterance for the speech emotion classification task. The baseline approach\nmodels the information from audio and text independently using two deep neural\nnetworks (DNNs). The outputs from both the DNNs are then fused for\nclassification. As opposed to using knowledge from both the modalities\nseparately, we propose a framework to exploit acoustic information in tandem\nwith lexical data. The proposed framework uses two bi-directional long\nshort-term memory (BLSTM) for obtaining hidden representations of the\nutterance. Furthermore, we propose an attention mechanism, referred to as the\nmulti-hop, which is trained to automatically infer the correlation between the\nmodalities. The multi-hop attention first computes the relevant segments of the\ntextual data corresponding to the audio signal. The relevant textual data is\nthen applied to attend parts of the audio signal. To evaluate the performance\nof the proposed system, experiments are performed in the IEMOCAP dataset.\nExperimental results show that the proposed technique outperforms the\nstate-of-the-art system by 6.5% relative improvement in terms of weighted\naccuracy.","url_abs":"http://arxiv.org/abs/1904.10788v2","url_pdf":"http://arxiv.org/pdf/1904.10788v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"190410788","repo_url":"https://github.com/raulsteleac/Speech_Emotion_Recognition","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}},{"paper_slug":"190410788","repo_url":"https://github.com/warnikchow/coaudiotext","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"emotion-classification","task_name":"Emotion Classification"},{"task_slug":"emotion-recognition","task_name":"Emotion Recognition"},{"task_slug":"speech-emotion-recognition","task_name":"Speech Emotion Recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1904.10788","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}