{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/wenlan-2-0-make-ai-imagine-via-a-multimodal","title":"Towards artificial general intelligence via a multimodal foundation model","arxiv_id":"2110.14378","date":"2021-10-27","proceeding":null,"authors":["Nanyi Fei","Zhiwu Lu","Yizhao Gao","Guoxing Yang","Yuqi Huo","Jingyuan Wen","Haoyu Lu","Ruihua Song","Xin Gao","Tao Xiang","Hao Sun","Ji-Rong Wen"],"abstract":"The fundamental goal of artificial intelligence (AI) is to mimic the core cognitive activities of human. Despite tremendous success in the AI research, most of existing methods have only single-cognitive ability. To overcome this limitation and take a solid step towards artificial general intelligence (AGI), we develop a foundation model pre-trained with huge multimodal data, which can be quickly adapted for various downstream cognitive tasks. To achieve this goal, we propose to pre-train our foundation model by self-supervised learning with weak semantic correlation data crawled from the Internet and show that promising results can be obtained on a wide range of downstream tasks. Particularly, with the developed model-interpretability tools, we demonstrate that strong imagination ability is now possessed by our foundation model. We believe that our work makes a transformative stride towards AGI, from our common practice of \"weak or narrow AI\" to that of \"strong or generalized AI\".","url_abs":"https://arxiv.org/abs/2110.14378v2","url_pdf":"https://arxiv.org/pdf/2110.14378v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"wenlan-2-0-make-ai-imagine-via-a-multimodal","repo_url":"https://github.com/neilfei/brivl-nmi","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok","spdx":"NOASSERTION"}}],"tasks":[{"task_slug":"image-classification","task_name":"Image Classification"},{"task_slug":"reading-comprehension","task_name":"Reading Comprehension"},{"task_slug":"self-supervised-learning","task_name":"Self-Supervised Learning"},{"task_slug":"visual-commonsense-reasoning","task_name":"Visual Commonsense Reasoning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2110.14378","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}