{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/flava-a-foundational-language-and-vision","title":"FLAVA: A Foundational Language And Vision Alignment Model","arxiv_id":"2112.04482","date":"2021-12-08","proceeding":"CVPR 2022 1","authors":["Amanpreet Singh","Ronghang Hu","Vedanuj Goswami","Guillaume Couairon","Wojciech Galuba","Marcus Rohrbach","Douwe Kiela"],"abstract":"State-of-the-art vision and vision-and-language models rely on large-scale visio-linguistic pretraining for obtaining good performance on a variety of downstream tasks. Generally, such models are often either cross-modal (contrastive) or multi-modal (with earlier fusion) but not both; and they often only target specific modalities or tasks. A promising direction would be to use a single holistic universal model, as a \"foundation\", that targets all modalities at once -- a true vision and language foundation model should be good at vision tasks, language tasks, and cross- and multi-modal vision and language tasks. We introduce FLAVA as such a model and demonstrate impressive performance on a wide range of 35 tasks spanning these target modalities.","url_abs":"https://arxiv.org/abs/2112.04482v3","url_pdf":"https://arxiv.org/pdf/2112.04482v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"flava-a-foundational-language-and-vision","repo_url":"https://github.com/apsdehal/flava-tutorials","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}},{"paper_slug":"flava-a-foundational-language-and-vision","repo_url":"https://github.com/facebookresearch/multimodal","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"flava-a-foundational-language-and-vision","repo_url":"https://github.com/social-ai-studio/matk","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"flava-a-foundational-language-and-vision","repo_url":"https://github.com/2024-MindSpore-1/Code2/tree/main/model-1/falcon","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null}],"tasks":[{"task_slug":"image-retrieval","task_name":"Image Retrieval"},{"task_slug":"image-to-text-retrieval","task_name":"Image-to-Text Retrieval"},{"task_slug":"visual-reasoning","task_name":"Visual Reasoning"},{"task_slug":"zero-shot-image-retrieval","task_name":"Zero-shot Image Retrieval"},{"task_slug":"zero-shot-text-retrieval","task_name":"Zero-shot Text Retrieval"},{"task_slug":"zero-shot-text-to-image-retrieval","task_name":"Zero-shot Text-to-Image Retrieval"}],"methods":[{"method_slug":"flava","method_name":"FLAVA"}],"datasets_introduced":[],"methods_introduced":[{"slug":"flava","name":"FLAVA","full_name":"FLAVA"}],"results":[{"leaderboard":"/sota/image-retrieval-on-coco","task":"Image Retrieval","dataset":"COCO (Common Objects in Context)","model":"FLAVA (zero-shot)","rank_in_archive_order":4,"of":6,"metrics":{"recall@1":"38.38","recall@5":"67.47"},"uses_additional_data":false},{"leaderboard":"/sota/image-retrieval-on-coco","task":"Image Retrieval","dataset":"COCO (Common Objects in Context)","model":"CLIP (zero-shot)","rank_in_archive_order":5,"of":6,"metrics":{"recall@1":"33.29","recall@5":"62.47"},"uses_additional_data":false},{"leaderboard":"/sota/image-to-text-retrieval-on-coco","task":"Image-to-Text Retrieval","dataset":"COCO (Common Objects in Context)","model":"FLAVA (ViT-B, zero-shot)","rank_in_archive_order":6,"of":9,"metrics":{"Recall@1":"42.74","Recall@5":"76.76"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=2112.04482","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}