{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/adversarial-multimodal-domain-transfer-for","title":"Adversarial Multimodal Domain Transfer for Video-Level Sentiment Analysis","arxiv_id":null,"date":"2022-05-11","proceeding":"IEEE Access 2022 5","authors":["Wang Yanan; Wu Jianming; Furumai Kazuaki; Wada Shinya; Kurihara Satoshi"],"abstract":"Video-level sentiment analysis is a challenging task and requires systems to obtain discriminative multimodal representations that can capture difference in sentiments across various modalities.\r\nHowever, due to diverse distributions of various modalities and the unified multimodal labels are not always\r\nadaptable to unimodal learning, the distance difference between unimodal representations increases, and\r\nprevents systems from learning discriminative multimodal representations. In this paper, to obtain more\r\ndiscriminative multimodal representations that can further improve systems’ performance, we propose a\r\nVAE-based adversarial multimodal domain transfer (VAE-AMDT) and jointly train it with a multi-attention\r\nmodule to reduce the distance difference between unimodal representations. We first perform variational\r\nautoencoder (VAE) to make visual, linguistic and acoustic representations follow a common distribution,\r\nand then introduce adversarial training to transfer all unimodal representations to a joint embedding space.\r\nAs a result, we fuse various modalities on this joint embedding space via the multi-attention module,\r\nwhich consists of self-attention, cross-attention and triple-attention for highlighting important sentimental\r\nrepresentations over time and modality. Our method improves F1-score of the state-of-the-art by 3.6%\r\non MOSI and 2.9% on MOSEI datasets, and prove its efficacy in obtaining discriminative multimodal\r\nrepresentations for video-level sentiment analysis.","url_abs":"https://ieeexplore.ieee.org/abstract/document/9772490","url_pdf":"https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=9772490","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"multimodal-sentiment-analysis","task_name":"Multimodal Sentiment Analysis"},{"task_slug":"sentiment-analysis","task_name":"Sentiment Analysis"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/multimodal-sentiment-analysis-on-cmu-mosi","task":"Multimodal Sentiment Analysis","dataset":"CMU-MOSI","model":"VAE-AMDT","rank_in_archive_order":6,"of":12,"metrics":{"Acc-2":"84.3","F1":"84.2","MAE":"0.716"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}