{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/end-to-end-multimodal-emotion-and-gender","title":"End-to-end Multimodal Emotion and Gender Recognition with Dynamic Joint Loss Weights","arxiv_id":"1809.00758","date":"2018-09-04","proceeding":null,"authors":["Myungsu Chae","Tae-Ho Kim","Young Hoon Shin","June-Woo Kim","Soo-Young Lee"],"abstract":"Multi-task learning is a method for improving the generalizability of\nmultiple tasks. In order to perform multiple classification tasks with one\nneural network model, the losses of each task should be combined. Previous\nstudies have mostly focused on multiple prediction tasks using joint loss with\nstatic weights for training models, choosing the weights between tasks without\nmaking sufficient considerations by setting them uniformly or empirically. In\nthis study, we propose a method to calculate joint loss using dynamic weights\nto improve the total performance, instead of the individual performance, of\ntasks. We apply this method to design an end-to-end multimodal emotion and\ngender recognition model using audio and video data. This approach provides\nproper weights for the loss of each task when the training process ends. In our\nexperiments, emotion and gender recognition with the proposed method yielded a\nlower joint loss, which is computed as the negative log-likelihood, than using\nstatic weights for joint loss. Moreover, our proposed model has better\ngeneralizability than other models. To the best of our knowledge, this research\nis the first to demonstrate the strength of using dynamic weights for joint\nloss for maximizing overall performance in emotion and gender recognition\ntasks.","url_abs":"http://arxiv.org/abs/1809.00758v3","url_pdf":"http://arxiv.org/pdf/1809.00758v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"end-to-end-multimodal-emotion-and-gender","repo_url":"https://github.com/MyungsuChae/IROS2018_ws","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"multi-task-learning","task_name":"Multi-Task Learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}