{"url":"/dataset/urdudoc","name":"UrduDoc","full_name":null,"description_markdown":"The **UrduDoc Dataset** is a benchmark dataset for Urdu text line detection in scanned documents. It is created as a byproduct of the **UTRSet-Real** dataset generation process. Comprising 478 diverse images collected from various sources such as books, documents, manuscripts, and newspapers, it offers a valuable resource for research in Urdu document analysis. It includes 358 pages for training and 120 pages for validation, featuring a wide range of styles, scales, and lighting conditions. It serves as a benchmark for evaluating printed Urdu text detection models, and the benchmark results of state-of-the-art models are provided. The Contour-Net model demonstrates the best performance in terms of h-mean.\r\n\r\nThe UrduDoc dataset is the first of its kind for printed Urdu text line detection and will advance research in the field. It will be made publicly available for non-commercial, academic, and research purposes upon request and execution of a no-cost license agreement. To request the dataset and for more information and details about the [UrduDoc ](https://paperswithcode.com/dataset/urdudoc), [UTRSet-Real](https://paperswithcode.com/dataset/utrset-real) & [UTRSet-Synth](https://paperswithcode.com/dataset/utrset-synth) datasets, please refer to the [Project Website](https://abdur75648.github.io/UTRNet/) of our paper [\"UTRNet: High-Resolution Urdu Text Recognition In Printed Documents\"](https://arxiv.org/abs/2306.15782)","description_withheld":null,"homepage":"https://abdur75648.github.io/UTRNet/","introduced_date":"2023-06-27","introduced_date_note":null,"introduced_by":{"paper":"/paper/utrnet-high-resolution-urdu-text-recognition","title":"UTRNet: High-Resolution Urdu Text Recognition In Printed Documents","first_author":"Abdur Rahman","url":null},"license":{"name":"CC BY-NC-ND","url":null},"modalities":[{"name":"Images","url":"/datasets/modality/images"},{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Scene Text Detection","url":"/task/scene-text-detection","datasets_with_task":"/datasets/task/scene-text-detection"},{"name":"Document Layout Analysis","url":"/task/document-layout-analysis","datasets_with_task":"/datasets/task/document-layout-analysis"},{"name":"Text Segmentation","url":"/task/text-segmentation","datasets_with_task":"/datasets/task/text-segmentation"},{"name":"Line Detection","url":"/task/line-detection","datasets_with_task":"/datasets/task/line-detection"},{"name":"Multi-Oriented Scene Text Detection","url":"/task/multi-oriented-scene-text-detection","datasets_with_task":"/datasets/task/multi-oriented-scene-text-detection"},{"name":"Curved Text Detection","url":"/task/curved-text-detection","datasets_with_task":"/datasets/task/curved-text-detection"},{"name":"Text-Line Extraction","url":"/task/text-line-extraction","datasets_with_task":"/datasets/task/text-line-extraction"},{"name":"Text Detection","url":"/task/text-detection","datasets_with_task":"/datasets/task/text-detection"}],"languages":[{"name":"Urdu","url":"/datasets/language/urdu"}],"variants":["UrduDoc"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/text-detection-on-urdudoc","task":"Text Detection","dataset_variant":"UrduDoc","rows":5,"metrics":["Precision","Recall"],"first_row_in_archive_order":{"model":"ContourNet [69]","paper":"/paper/utrnet-high-resolution-urdu-text-recognition","metrics":{"Precision":"86.99","Recall":"88.68"},"code_links":[{"title":"abdur75648/UTRNet-High-Resolution-Urdu-Text-Recognition","url":"https://github.com/abdur75648/UTRNet-High-Resolution-Urdu-Text-Recognition"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/utrnet-high-resolution-urdu-text-recognition","title":"UTRNet: High-Resolution Urdu Text Recognition In Printed Documents","date":"2023-06-27","rows_on_this_dataset":5,"code_links":1,"syntology":null}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}