{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/investigation-of-multimodal-features","title":"Investigation of Multimodal Features, Classifiers and Fusion Methods for Emotion Recognition","arxiv_id":"1809.06225","date":"2018-09-13","proceeding":null,"authors":["Zheng Lian","Ya Li","Jian-Hua Tao","Jian Huang"],"abstract":"Automatic emotion recognition is a challenging task. In this paper, we\npresent our effort for the audio-video based sub-challenge of the Emotion\nRecognition in the Wild (EmotiW) 2018 challenge, which requires participants to\nassign a single emotion label to the video clip from the six universal emotions\n(Anger, Disgust, Fear, Happiness, Sad and Surprise) and Neutral. The proposed\nmultimodal emotion recognition system takes audio, video and text information\ninto account. Except for handcraft features, we also extract bottleneck\nfeatures from deep neutral networks (DNNs) via transfer learning. Both temporal\nclassifiers and non-temporal classifiers are evaluated to obtain the best\nunimodal emotion classification result. Then possibilities are extracted and\npassed into the Beam Search Fusion (BS-Fusion). We test our method in the\nEmotiW 2018 challenge and we gain promising results. Compared with the baseline\nsystem, there is a significant improvement. We achieve 60.34% accuracy on the\ntesting dataset, which is only 1.5% lower than the winner. It shows that our\nmethod is very competitive.","url_abs":"http://arxiv.org/abs/1809.06225v1","url_pdf":"http://arxiv.org/pdf/1809.06225v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"investigation-of-multimodal-features","repo_url":"https://github.com/zeroQiaoba/EmotiW2018","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"emotion-classification","task_name":"Emotion Classification"},{"task_slug":"emotion-recognition","task_name":"Emotion Recognition"},{"task_slug":"multimodal-emotion-recognition","task_name":"Multimodal Emotion Recognition"},{"task_slug":"transfer-learning","task_name":"Transfer Learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}