{"url":"/dataset/arabic-img2md","name":"arabic-img2md","full_name":"Arabic Img2MD","description_markdown":"Click to add a brief description of the dataset (Markdown and LaTe# Arabic Img2MD\r\n\r\n## Dataset Summary\r\n\r\nThe `arabic-img2md` dataset consists of **15,000 examples** of PDF pages paired with their Markdown counterparts. The dataset is split into:\r\n- **Train:** 13,700 examples\r\n- **Test:** 1,520 examples\r\n\r\nThis dataset was created as part of the open-source research project **Arabic Nougat** to enable OCR and Markdown extraction from Arabic documents. It contains mostly **Arabic text** but also includes examples with **English text**.\r\nX enabled).\r\n\r\nProvide:\r\n\r\n* a high-level explanation of the dataset characteristics\r\n* explain motivations and summary of its content\r\n* potential use cases of the dataset","description_withheld":null,"homepage":"https://huggingface.co/datasets/MohamedRashad/arabic-img2md","introduced_date":"2024-11-19","introduced_date_note":null,"introduced_by":{"paper":"/paper/arabic-nougat-fine-tuning-vision-transformers","title":"Arabic-Nougat: Fine-Tuning Vision Transformers for Arabic OCR and Markdown Extraction","first_author":"Mohamed Rashad","url":null},"license":null,"modalities":[],"tasks":[{"name":"Optical Character Recognition (OCR)","url":"/task/optical-character-recognition","datasets_with_task":"/datasets/task/optical-character-recognition"}],"languages":[{"name":"Arabic","url":"/datasets/language/arabic"}],"variants":["arabic-img2md"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}