{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/attention-based-fully-convolutional-network","title":"Attention Based Fully Convolutional Network for Speech Emotion Recognition","arxiv_id":"1806.01506","date":"2018-06-05","proceeding":null,"authors":["Yuanyuan Zhang","Jun Du","Zi-Rui Wang","Jianshu Zhang"],"abstract":"Speech emotion recognition is a challenging task for three main reasons: 1)\nhuman emotion is abstract, which means it is hard to distinguish; 2) in\ngeneral, human emotion can only be detected in some specific moments during a\nlong utterance; 3) speech data with emotional labeling is usually limited. In\nthis paper, we present a novel attention based fully convolutional network for\nspeech emotion recognition. We employ fully convolutional network as it is able\nto handle variable-length speech, free of the demand of segmentation to keep\ncritical information not lost. The proposed attention mechanism can make our\nmodel be aware of which time-frequency region of speech spectrogram is more\nemotion-relevant. Considering limited data, the transfer learning is also\nadapted to improve the accuracy. Especially, it's interesting to observe\nobvious improvement obtained with natural scene image based pre-trained model.\nValidated on the publicly available IEMOCAP corpus, the proposed model\noutperformed the state-of-the-art methods with a weighted accuracy of 70.4% and\nan unweighted accuracy of 63.9% respectively.","url_abs":"http://arxiv.org/abs/1806.01506v2","url_pdf":"http://arxiv.org/pdf/1806.01506v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"attention-based-fully-convolutional-network","repo_url":"https://github.com/aris-ai/Audio-and-text-based-emotion-recognition","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"emotion-recognition","task_name":"Emotion Recognition"},{"task_slug":"speech-emotion-recognition","task_name":"Speech Emotion Recognition"},{"task_slug":"transfer-learning","task_name":"Transfer Learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}