Datasets › TexBiG
TexBiG (Text-Bild-Gefüge)
TexBiG (from the German Text-Bild-Gefüge, meaning Text-Image-Structure) is a document layout analysis dataset for historical documents in the late 19th and early 20th century. The dataset provides instance segmentation (bounding boxes and polygons/masks) annotations for 19 different classes with more then 52.000 instances.
The added test images can be used to make submission on the leaderboard on EvalAI: https://eval.ai/web/challenges/challenge-page/2078/overview
Dataset link: https://doi.org/10.5281/zenodo.8347059
Dataset (only train): https://www.kaggle.com/datasets/davidtschirschwitz/texbig-v2-0-train-val
Dataset (only test image): https://www.kaggle.com/datasets/davidtschirschwitz/texbig-v2-0-test
Please use TexBiG 2023 (which is v2.0 of the dataset) for testing model performance. The test dataset from the 2022 version (v1.0) are included as training data in v2.0
Each image of the dataset was annotated at least by two different annotators.
Benchmarks archive 2025-07-28
All 4 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.
| First row (archive order) | Paper | Code | ||||
|---|---|---|---|---|---|---|
| Instance Segmentation | TexBiG 2022 test | VSR (Vison, Semantics and Relation Model) mAP@0.5:0.95:0.05 65.8 | A Dataset for Analysing Complex Document Layouts in the... | Madave94/kalphacv +1 | 1 | Compare |
| Instance Segmentation | TexBiG 2023 test | DetectoRS + LAEM mAP@0.5:0.95:0.05 44.06 | Drawing the Same Bounding Box Twice? Coping Noisy... | madave94/gtiod | 1 | Compare |
| Object Detection | TexBiG 2022 test | VSR (Vison, Semantics and Relation Model) mAP@0.5:0.95:0.05 75.9 | A Dataset for Analysing Complex Document Layouts in the... | Madave94/kalphacv +1 | 1 | Compare |
| Object Detection | TexBiG 2023 test | DetectoRS + LAEM mAP@0.5:0.95:0.05 49.89 | Drawing the Same Bounding Box Twice? Coping Noisy... | madave94/gtiod | 1 | Compare |
Papers archive 2025-07-28
2 shown of 2 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 4. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.
| Date | Samples run Syntology | |||
|---|---|---|---|---|
| Drawing the Same Bounding Box Twice? Coping Noisy Annotations in Object Detection with Repeated Labels | 1 | 2 | 18 Sep 2023 | not harvested |
| A Dataset for Analysing Complex Document Layouts in the Digital Humanities and Its Evaluation with Krippendorff’s Alpha | 2 | 2 | 20 Sep 2022 | not harvested |
Dataset loaders archive 2025-07-28
No loader listed in the archive.
Tasks archive 2025-07-28
License archive 2025-07-28
Creative Commons Attribution 4.0 International
Modalities archive 2025-07-28
Languages archive 2025-07-28
Variants archive 2025-07-28
- TexBiG
- TexBiG 2023 test
- TexBiG 2022
- TexBiG 2022 test
4 variant names, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections