{"url":"/dataset/urdu-msmarco","name":"Urdu MsMarco","full_name":null,"description_markdown":"This dataset is the translation of the MS-marco dataset, marking it the first large-scale urdu IR dataset.\r\n\r\nDataset Details:\r\nThe MS MARCO dataset is formed by a collection of 8.8M passages, approximately 530k queries, and at least one relevant passage per query, which were selected by humans. The development set of MS MARCO comprises more than 100k queries. However, a smaller set of 6,980 queries is used for evaluation in most published works.\r\n\r\nThe triples files (triples.train.small.urdu.tsv , named as triple files part aa to ae) is around 47 GB and is split into 5 parts. Download them and combine them before you start working.\r\n\r\nThis dataset is created using the IndicTrans2 translation model.","description_withheld":null,"homepage":"https://huggingface.co/datasets/Mavkif/urdu-msmarco-dataset","introduced_date":"2024-12-17","introduced_date_note":null,"introduced_by":{"paper":"/paper/enabling-low-resource-language-retrieval","title":"Enabling Low-Resource Language Retrieval: Establishing Baselines for Urdu MS MARCO","first_author":"Umer Butt","url":null},"license":{"name":"CC BY-SA","url":null},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Information Retrieval","url":"/task/information-retrieval","datasets_with_task":"/datasets/task/information-retrieval"}],"languages":[{"name":"Urdu","url":"/datasets/language/urdu"}],"variants":["Urdu MsMarco"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}