{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/wit-wikipedia-based-image-text-dataset-for","title":"WIT: Wikipedia-based Image Text Dataset for Multimodal Multilingual Machine Learning","arxiv_id":"2103.01913","date":"2021-03-02","proceeding":null,"authors":["Krishna Srinivasan","Karthik Raman","Jiecao Chen","Michael Bendersky","Marc Najork"],"abstract":"The milestone improvements brought about by deep representation learning and pre-training techniques have led to large performance gains across downstream NLP, IR and Vision tasks. Multimodal modeling techniques aim to leverage large high-quality visio-linguistic datasets for learning complementary information (across image and text modalities). In this paper, we introduce the Wikipedia-based Image Text (WIT) Dataset (https://github.com/google-research-datasets/wit) to better facilitate multimodal, multilingual learning. WIT is composed of a curated set of 37.6 million entity rich image-text examples with 11.5 million unique images across 108 Wikipedia languages. Its size enables WIT to be used as a pretraining dataset for multimodal models, as we show when applied to downstream tasks such as image-text retrieval. WIT has four main and unique advantages. First, WIT is the largest multimodal dataset by the number of image-text examples by 3x (at the time of writing). Second, WIT is massively multilingual (first of its kind) with coverage over 100+ languages (each of which has at least 12K examples) and provides cross-lingual texts for many images. Third, WIT represents a more diverse set of concepts and real world entities relative to what previous datasets cover. Lastly, WIT provides a very challenging real-world test set, as we empirically illustrate using an image-text retrieval task as an example.","url_abs":"https://arxiv.org/abs/2103.01913v2","url_pdf":"https://arxiv.org/pdf/2103.01913v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"wit-wikipedia-based-image-text-dataset-for","repo_url":"https://github.com/google-research-datasets/wit","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"NOASSERTION"}},{"paper_slug":"wit-wikipedia-based-image-text-dataset-for","repo_url":"https://github.com/clip-italian/clip-italian","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"jax","reach":{"status":"unanswered"}},{"paper_slug":"wit-wikipedia-based-image-text-dataset-for","repo_url":"https://github.com/paullerner/viquae","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"NOASSERTION"}}],"tasks":[{"task_slug":"machine-learning","task_name":"BIG-bench Machine Learning"},{"task_slug":"image-retrieval","task_name":"Image Retrieval"},{"task_slug":"image-text-retrieval","task_name":"Image-text Retrieval"},{"task_slug":"representation-learning","task_name":"Representation Learning"},{"task_slug":"retrieval","task_name":"Retrieval"},{"task_slug":"text-retrieval","task_name":"Text Retrieval"}],"methods":[],"datasets_introduced":[{"slug":"wit","name":"WIT","full_name":"Wikipedia-based Image Text"}],"methods_introduced":[],"results":[{"leaderboard":"/sota/image-retrieval-on-wit","task":"Image Retrieval","dataset":"WIT","model":"WIT-ALL","rank_in_archive_order":1,"of":2,"metrics":{"R@1":"0.346","R@5":"0.642"},"uses_additional_data":false},{"leaderboard":"/sota/image-retrieval-on-wit","task":"Image Retrieval","dataset":"WIT","model":"CC (Conceptual Captions)","rank_in_archive_order":2,"of":2,"metrics":{"R@1":"0.048","R@5":"0.122"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2103.01913","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}