{"url":"/dataset/sosd","name":"SOSD","full_name":"Searching on Sorted Data","description_markdown":"SOSD is a collection of dataset to benchmark the lookup performance of learned indexes.\r\n\r\nSOSD currently includes eight different datasets. Each dataset consists of 200 million 64-bit unsigned integers (keys) with very few duplicates (if at all):\r\n`amzn` represents book sale popularity data.\r\n`face` is an upsampled version of a Facebook user ID dataset.\r\n`logn` and `norm` are lognormal (0, 2) and normal distributions, respectively.\r\n`osmc` is uniformly sampled OpenStreetMap locations represented as Google S2 CellIds.\r\n`uden` is dense integers.\r\n`uspr` is uniformly distributed sparse integers.\r\n`wiki` is Wikipedia article edit timestamps.\r\n\r\nIn addition, there are 32-bit versions of all datasets (except `osmc` and `wiki`) with similar CDFs. We use different parameters, (0, 1), for logn in the 32-bit case to reduce the number of duplicates.","description_withheld":null,"homepage":"https://dataverse.harvard.edu/dataset.xhtml?persistentId=doi:10.7910/DVN/JGVF9A","introduced_date":"2019-08-01","introduced_date_note":null,"introduced_by":{"paper":"/paper/sosd-a-benchmark-for-learned-indexes","title":"SOSD: A Benchmark for Learned Indexes","first_author":"Andreas Kipf","url":null},"license":{"name":"CC0","url":"https://dataverse.org/best-practices/dataverse-community-norms"},"modalities":[],"tasks":[],"languages":[],"variants":["SOSD"],"data_loaders":[],"num_papers_in_archive":17,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-25T09:33:49+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}