Datasets › UrduDoc

UrduDoc

Introduced by Abdur Rahman et al. in UTRNet: High-Resolution Urdu Text Recognition In Printed Documents27 Jun 2023 archive 2025-07-28

The UrduDoc Dataset is a benchmark dataset for Urdu text line detection in scanned documents. It is created as a byproduct of the UTRSet-Real dataset generation process. Comprising 478 diverse images collected from various sources such as books, documents, manuscripts, and newspapers, it offers a valuable resource for research in Urdu document analysis. It includes 358 pages for training and 120 pages for validation, featuring a wide range of styles, scales, and lighting conditions. It serves as a benchmark for evaluating printed Urdu text detection models, and the benchmark results of state-of-the-art models are provided. The Contour-Net model demonstrates the best performance in terms of h-mean.

The UrduDoc dataset is the first of its kind for printed Urdu text line detection and will advance research in the field. It will be made publicly available for non-commercial, academic, and research purposes upon request and execution of a no-cost license agreement. To request the dataset and for more information and details about the UrduDoc , UTRSet-Real & UTRSet-Synth datasets, please refer to the Project Website of our paper "UTRNet: High-Resolution Urdu Text Recognition In Printed Documents"

Benchmarks archive 2025-07-28

All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Text Detection UrduDoc ContourNet [69] Precision 86.99 UTRNet: High-Resolution Urdu Text Recognition In Printed... abdur75648/UTRNet-High-Resolution-Urdu-Text-Recognition 5 Compare

Papers archive 2025-07-28

1 shown of 1 paper with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 1. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
UTRNet: High-Resolution Urdu Text Recognition In Printed Documents 1 5 27 Jun 2023 not harvested

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

CC BY-NC-ND

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • UrduDoc

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections