{"url":"/dataset/openstreetview-5m","name":"OpenStreetView-5M","full_name":null,"description_markdown":"OpenStreetView-5M establishes a new open benchmark for geolocation by providing a large, open, and clean dataset. As detailed below, OpenStreetView-5M improves upon several limitations of current geolocation datasets.\r\n\r\nDeep neural networks have historically been selected over other machine learning methods because they benefit from larger amounts of data. OSV-5M consists of 4,894,685 training and 210,122 test images, with a height of $512$ pixels and an average width of $792\\pm127$  pixels.\r\n\r\n Many geolocation datasets are restricted to a few cities or are significantly biased towards the Western world . In contrast, OpenStreetView-5M images are uniformly sampled on the globe, covering 70k cities and 225 countries and territories. The distribution of test images across countries has a normalized entropy of $0.78$~\\cite[Eq. 19]{wilcox1967indices}, suggesting high diversity. Our train set has a normalized entropy of $0.67$, which is comparable to the entropy of the distribution of the countries' area ($0.71$).\r\n\r\n OpenStreetView-5M is based on the crowd-sourced street view images of Mapillary  which follow the CC-BY-SA license: free of use with attribution.\r\n\r\n\r\nWe estimate through manual inspection of 4500 images that 96.1\\% (±0.57\\%) of the images in the OpenStreetView-5M dataset are localizable, with a 95\\% confidence level. Among the weakly or non-localizable images, $70$\\% ($2.7$\\% total) are low-quality: under- or over-exposed, blurry, or rotated; $30$\\% ($1.2$\\% total) are poorly framed, indoor, or in tunnels.\r\n\r\n Without carefully enforcing the spatial separation between train and test images, geolocation can reduce to place-recognition. As our goal is to assess the capacity of models to learn robust geographical representations, we ensure that no image in the OSV-5M training set lies within a $1$km radius of any image in the test set.\r\n\r\n\r\nStreet-view images are typically acquired by a limited number of camera sensors mounted on the top or front of a small fleet of vehicles assigned to a given region. This correlation between location, cars, and sensors can be exploited to simplify the geolocation task. Notoriously, players of the web-based geolocation game GeoGuessr  can locate images from Ghana by spotting a piece of duct tape placed on the corner of the roof rack of the Google Street View car . OpenStreetView-5M tries to avoid this pitfall by ensuring that no image sequence (a continuous series of images acquired by the same user) appears in both training and test sets. While this might not prevent images taken with the same vehicle on different days from being in both sets, it limits such occurrences.\r\n\r\nRich metadata beyond geographical coordinates can improve the robustness and versatility of geolocation models. Each image in our dataset is associated with four tiers of administrative data: country, region (\\eg, state), area (\\eg, county), and the nearest city. Note that areas are not defined for one-third of the dataset.\r\n\r\nWe also associate each image with a set of additional information: land cover, climate, soil type, the driving side, and distance to the sea where the image was taken.","description_withheld":null,"homepage":"https://huggingface.co/datasets/osv5m/osv5m","introduced_date":"2024-04-29","introduced_date_note":null,"introduced_by":{"paper":"/paper/openstreetview-5m-the-many-roads-to-global","title":"OpenStreetView-5M: The Many Roads to Global Visual Geolocation","first_author":"Guillaume Astruc","url":null},"license":{"name":"CC-BY-SA","url":null},"modalities":[],"tasks":[{"name":"Photo geolocation estimation","url":"/task/photo-geolocation-estimation","datasets_with_task":"/datasets/task/photo-geolocation-estimation"}],"languages":[],"variants":["OpenStreetView-5M"],"data_loaders":[],"num_papers_in_archive":2,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/photo-geolocation-estimation-on","task":"Photo geolocation estimation","dataset_variant":"OpenStreetView-5M","rows":2,"metrics":["Geoscore"],"first_row_in_archive_order":{"model":"Plonk","paper":"/paper/around-the-world-in-80-timesteps-a-generative","metrics":{"Geoscore":"3767"},"code_links":[{"title":"nicolas-dufour/plonk","url":"https://github.com/nicolas-dufour/plonk"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/around-the-world-in-80-timesteps-a-generative","title":"Around the World in 80 Timesteps: A Generative Approach to Global Visual Geolocation","date":"2024-12-09","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":7,"samples_ran":0,"samples_unverified":7,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/openstreetview-5m-the-many-roads-to-global","title":"OpenStreetView-5M: The Many Roads to Global Visual Geolocation","date":"2024-04-29","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":1,"samples_ran":0,"samples_unverified":1,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":2,"samples_harvested":8,"samples_ran":0,"samples_unverified":8,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":2,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}