Papers › SUT: a new multi-purpose synthetic dataset for Farsi document image analysis
SUT: a new multi-purpose synthetic dataset for Farsi document image analysis
Elham Shabaninia, Fatemeh sadat Eslami, Ali Afkari Fahandari, Hossein Nezamabadi-pour
This paper introduces a new large-scale dataset for Farsi document images, named SUT, which aims to tackle the challenges associated with obtaining diverse and substantial ground-truth data for supervised models in document image analysis (DIA) tasks, such as document image classification, text detection and recognition, and information retrieval. The dataset comprises 62,453 images that have been categorized into 21 distinct classes, including identity documents featuring synthetically generated personal information superimposed on various backgrounds. The dataset also includes corresponding files with labeling information for the images. The ground-truth data is organized in CSV files containing compiled image file paths and associated information about the embedded data. To demonstrate the efficacy of the SUT dataset in DIA tasks, it was utilized for document classification (achieving an accuracy of 86% using a convolutional neural network) and OCR (achieving a CER of 0.083 and 0.072 using Tesseract and EasyOCR engines, respectively). The SUT dataset represents a valuable resource for researchers who are interested in developing and evaluating supervised models in Farsi document image analysis.
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Document Image Classification | SUT | CNN | Accuracy | 86% | #1 of 1 | Archive leaderboard | report |
| Optical Character Recognition (OCR) | SUT | Tesseract | Character Error Rate (CER) | 0.083 | #1 of 2 | Archive leaderboard | report |
| Optical Character Recognition (OCR) | SUT | EasyOCR | Character Error Rate (CER) | 0.072 | #2 of 2 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections