Datasets › BN-HTRd
BN-HTRd (BN-HTRd: A Benchmark Dataset for Document Level Offline Bangla Handwritten Text Recognition (HTR))
We introduce a new Dataset (BN-HTRd) for offline Handwritten Text Recognition (HTR) from images of Bangla scripts comprising words, lines, and document-level annotations. The BN-HTRd dataset is based on the BBC Bangla News corpus - which acted as ground truth texts for the handwritings. Our dataset contains a total of 786 full-page images collected from 150 different writers. With a staggering 1,08,181 instances of handwritten words, distributed over 14,383 lines and 23,115 unique words, this is currently the 'largest and most comprehensive dataset' in this field. We also provided the bounding box annotations (YOLO format) for the segmentation of words/lines and the ground truth annotations for full-text, along with the segmented images and their positions. The contents of our dataset came from a diverse news category, and annotators of different ages, genders, and backgrounds, having variability in writing styles. The BN-HTRd dataset can be adopted as a basis for various handwriting classification tasks such as end-to-end document recognition, word-spotting, word/line segmentation, and so on.
The statistics of the original dataset are given below:
- Number of writers = 150
- Total number of images = 786
- Total number of lines = 14,383
- Total number of words = 1,08,181
- Total number of unique words = 23,115
- Total number of punctuation = 7,446
- Total number of characters = 5,74,203
From v3.0 onwards, we are also providing automatic bounding box annotations (YOLO format) of 805 document images containing words/lines. The statistics of the automatic annotations are given below:
- Number of writers = 87
- Total number of images = 805
- Total number of lines = 14,836
- Total number of words = 1,06,135
Benchmarks archive 2025-07-28
All 2 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.
| First row (archive order) | Paper | Code | ||||
|---|---|---|---|---|---|---|
| Handwritten Line Segmentation | BN-HTRd | BN-DRISHTI Line Segmentation F-Score 0.9997 | BN-DRISHTI: Bangla Document Recognition through... | crusnic-corp/BN-DRISHTI | 2 | Compare |
| Handwritten Word Segmentation | BN-HTRd | BN-DRISHTI Word Segmentation F-Score 0.98 | BN-DRISHTI: Bangla Document Recognition through... | crusnic-corp/BN-DRISHTI | 1 | Compare |
Papers archive 2025-07-28
2 shown of 2 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 2. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.
| Date | Samples run Syntology | |||
|---|---|---|---|---|
| BN-DRISHTI: Bangla Document Recognition through Instance-level Segmentation of Handwritten Text Images | 1 | 2 | 31 May 2023 | not harvested |
| BN-HTRd: A Benchmark Dataset for Document Level Offline Bangla Handwritten Text Recognition (HTR) and Line Segmentation | 1 | 1 | 29 May 2022 | not harvested |
Dataset loaders archive 2025-07-28
No loader listed in the archive.
Tasks archive 2025-07-28
License archive 2025-07-28
Modalities archive 2025-07-28
No modality tagged.
Languages archive 2025-07-28
Variants archive 2025-07-28
- BN-HTRd
1 variant name, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections