Datasets › Urdu MsMarco

Urdu MsMarco

Introduced by Umer Butt et al. in Enabling Low-Resource Language Retrieval: Establishing Baselines for Urdu MS MARCO17 Dec 2024 archive 2025-07-28

This dataset is the translation of the MS-marco dataset, marking it the first large-scale urdu IR dataset.

Dataset Details: The MS MARCO dataset is formed by a collection of 8.8M passages, approximately 530k queries, and at least one relevant passage per query, which were selected by humans. The development set of MS MARCO comprises more than 100k queries. However, a smaller set of 6,980 queries is used for evaluation in most published works.

The triples files (triples.train.small.urdu.tsv , named as triple files part aa to ae) is around 47 GB and is split into 5 parts. Download them and combine them before you start working.

This dataset is created using the IndicTrans2 translation model.

Benchmarks archive 2025-07-28

No leaderboard in the archive resolves to this dataset.

Papers archive 2025-07-28

No paper in the archive has a leaderboard row on this dataset; the archive counts 1 paper for it but never published that list.

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

CC BY-SA

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • Urdu MsMarco

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections