Datasets › Hindi Text Image Dataset | Hindi in the wild
Hindi Text Image Dataset | Hindi in the wild (datacluster.ai)
This dataset is an extremely challenging set of over 5000+ original Hindi text images captured and crowdsourced from over 700+ urban and rural areas, where each image is manually reviewed and verified by computer vision professionals at DataclusterLabs.
Dataset Features
- Dataset size : 5000+
- Captured by : Over 700+ crowdsource contributors
- Resolution : 99% images HD and above (1920x1080 and above)
- Location : Captured with 400+ cities accross India
- Diversity : Various lighting conditions like day, night, varied distances, view points etc.
- Device used : Captured using mobile phones in 2021-2022
- Usage : Hindi text detection, Hindi NLP, Text recognition, etc.
Available Annotation formats
COCO, YOLO, PASCAL-VOC, Tf-Record
To download full datasets or to submit a request for your dataset needs, please ping us at sales@datacluster.ai Visit www.datacluster.ai to know more.
Note:
All the images are manually captured and verified by a large contributor base on DataCluster platform
Benchmarks archive 2025-07-28
No leaderboard in the archive resolves to this dataset.
Papers archive 2025-07-28
No paper in the archive has a leaderboard row on this dataset.
Dataset loaders archive 2025-07-28
No loader listed in the archive.
Tasks archive 2025-07-28
License archive 2025-07-28
No licence recorded in the archive. Absence here is not a statement about the dataset's terms.
Modalities archive 2025-07-28
Languages archive 2025-07-28
No language tagged.
Variants archive 2025-07-28
- Hindi Text Image Dataset | Hindi in the wild
1 variant name, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections