Datasets › Hindi Text Image Dataset | Hindi in the wild

Hindi Text Image Dataset | Hindi in the wild (datacluster.ai)

archive 2025-07-28

This dataset is an extremely challenging set of over 5000+ original Hindi text images captured and crowdsourced from over 700+ urban and rural areas, where each image is manually reviewed and verified by computer vision professionals at DataclusterLabs.

Dataset Features
  • Dataset size : 5000+
  • Captured by : Over 700+ crowdsource contributors
  • Resolution : 99% images HD and above (1920x1080 and above)
  • Location : Captured with 400+ cities accross India
  • Diversity : Various lighting conditions like day, night, varied distances, view points etc.
  • Device used : Captured using mobile phones in 2021-2022
  • Usage : Hindi text detection, Hindi NLP, Text recognition, etc.
Available Annotation formats

COCO, YOLO, PASCAL-VOC, Tf-Record

To download full datasets or to submit a request for your dataset needs, please ping us at sales@datacluster.ai Visit www.datacluster.ai to know more.

Note:
All the images are manually captured and verified by a large contributor base on DataCluster platform

Benchmarks archive 2025-07-28

No leaderboard in the archive resolves to this dataset.

Papers archive 2025-07-28

No paper in the archive has a leaderboard row on this dataset.

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • Hindi Text Image Dataset | Hindi in the wild

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections