{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/deep-multimodality-model-for-multi-task-multi","title":"Deep Multimodality Model for Multi-task Multi-view Learning","arxiv_id":"1901.08723","date":"2019-01-25","proceeding":null,"authors":["Lecheng Zheng","Yu Cheng","Jingrui He"],"abstract":"Many real-world problems exhibit the coexistence of multiple types of\nheterogeneity, such as view heterogeneity (i.e., multi-view property) and task\nheterogeneity (i.e., multi-task property). For example, in an image\nclassification problem containing multiple poses of the same object, each pose\ncan be considered as one view, and the detection of each type of object can be\ntreated as one task. Furthermore, in some problems, the data type of multiple\nviews might be different. In a web classification problem, for instance, we\nmight be provided an image and text mixed data set, where the web pages are\ncharacterized by both images and texts. A common strategy to solve this kind of\nproblem is to leverage the consistency of views and the relatedness of tasks to\nbuild the prediction model. In the context of deep neural network, multi-task\nrelatedness is usually realized by grouping tasks at each layer, while\nmulti-view consistency is usually enforced by finding the maximal correlation\ncoefficient between views. However, there is no existing deep learning\nalgorithm that jointly models task and view dual heterogeneity, particularly\nfor a data set with multiple modalities (text and image mixed data set or text\nand video mixed data set, etc.). In this paper, we bridge this gap by proposing\na deep multi-task multi-view learning framework that learns a deep\nrepresentation for such dual-heterogeneity problems. Empirical studies on\nmultiple real-world data sets demonstrate the effectiveness of our proposed\nDeep-MTMV algorithm.","url_abs":"http://arxiv.org/abs/1901.08723v1","url_pdf":"http://arxiv.org/pdf/1901.08723v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"deep-multimodality-model-for-multi-task-multi","repo_url":"https://github.com/Leo02016/DeepMTMV","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"image-classification","task_name":"Image Classification"},{"task_slug":"multi-view-learning","task_name":"MULTI-VIEW LEARNING"},{"task_slug":"image-classification","task_name":"image-classification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}