{"url":"/dataset/wikiscenes","name":"WikiScenes","full_name":null,"description_markdown":"The **WikiScenes** dataset consists of paired images and language descriptions capturing world landmarks and cultural sites, with associated 3D models and camera poses. WikiScenes is derived from the massive public catalog of freely-licensed crowdsourced data in the Wikimedia Commons project, which contains a large variety of images with captions and other metadata. \r\n\r\nThe dataset contains two forms of textual descriptions for each image: (1) Captions associated with images, describing the image using free-form language, and (2) The WikiCategory hierarchy obtained according to the hierarchy of WikiCategories associated with each image (see the examples in the image below). Overall, WikiScenes contains approximately 63K images with textual descriptions.","description_withheld":null,"homepage":"https://www.cs.cornell.edu/projects/babel/","introduced_date":"2021-08-12","introduced_date_note":null,"introduced_by":{"paper":"/paper/towers-of-babel-combining-images-language-and","title":"Towers of Babel: Combining Images, Language, and 3D Geometry for Learning Multimodal Vision","first_author":"Xiaoshi Wu","url":null},"license":null,"modalities":[{"name":"Images","url":"/datasets/modality/images"},{"name":"Texts","url":"/datasets/modality/texts"},{"name":"3D","url":"/datasets/modality/3d"}],"tasks":[{"name":"Image Captioning","url":"/task/image-captioning","datasets_with_task":"/datasets/task/image-captioning"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["WikiScenes"],"data_loaders":[],"num_papers_in_archive":3,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-25T09:33:49+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}