{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/joint-multimodal-learning-with-deep","title":"Joint Multimodal Learning with Deep Generative Models","arxiv_id":"1611.01891","date":"2016-11-07","proceeding":null,"authors":["Masahiro Suzuki","Kotaro Nakayama","Yutaka Matsuo"],"abstract":"We investigate deep generative models that can exchange multiple modalities\nbi-directionally, e.g., generating images from corresponding texts and vice\nversa. Recently, some studies handle multiple modalities on deep generative\nmodels, such as variational autoencoders (VAEs). However, these models\ntypically assume that modalities are forced to have a conditioned relation,\ni.e., we can only generate modalities in one direction. To achieve our\nobjective, we should extract a joint representation that captures high-level\nconcepts among all modalities and through which we can exchange them\nbi-directionally. As described herein, we propose a joint multimodal\nvariational autoencoder (JMVAE), in which all modalities are independently\nconditioned on joint representation. In other words, it models a joint\ndistribution of modalities. Furthermore, to be able to generate missing\nmodalities from the remaining modalities properly, we develop an additional\nmethod, JMVAE-kl, that is trained by reducing the divergence between JMVAE's\nencoder and prepared networks of respective modalities. Our experiments show\nthat our proposed method can obtain appropriate joint representation from\nmultiple modalities and that it can generate and reconstruct them more properly\nthan conventional VAEs. We further demonstrate that JMVAE can generate multiple\nmodalities bi-directionally.","url_abs":"http://arxiv.org/abs/1611.01891v1","url_pdf":"http://arxiv.org/pdf/1611.01891v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"joint-multimodal-learning-with-deep","repo_url":"https://github.com/masa-su/Tars","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"joint-multimodal-learning-with-deep","repo_url":"https://github.com/masa-su/jmvae","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1611.01891","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}