Datasets › SIMARA
SIMARA (SIMARA: a database for key-value information extraction from full-page handwritten documents)
Description
We propose a new database for information extraction from historical handwritten documents. The corpus includes 5,393 finding aids from six different series, dating from the 18th-20th centuries. Finding aids are handwritten documents that contain metadata describing older archives. They are stored in the National Archives of France and are used by archivists to identify and find archival documents.
Each document is annotated at page-level, and contains seven fields to retrieve. The localization of each field is not available in such a way that this dataset encourages research on segmentation-free systems for information extraction.
The dataset is available at https://zenodo.org/record/7868059
Details for each series and entity type
| Series | Train | Validation | Test | Total (%) |
|---|---|---|---|---|
| E series | 322 | 64 | 79 | 8.6 |
| L series | 38 | 8 | 4 | 0.9 |
| M series | 128 | 21 | 27 | 3.3 |
| X1a series | 2209 | 491 | 469 | 58.8 |
| Y series | 940 | 205 | 196 | 24.9 |
| Douët s'Arcq series | 141 | 22 | 29 | 3.5 |
| Total | 3778 | 811 | 804 | 100 |
| Entities | Train | Validation | Test | Total (%) |
|---|---|---|---|---|
| date | 8406 | 1814 | 1799 | 10.4 |
| title | 35531 | 7495 | 8173 | 44.5 |
| serie | 3168 | 664 | 676 | 3.9 |
| analysis | 25988 | 5130 | 5602 | 31.9 |
| volume_number | 3913 | 808 | 813 | 4.8 |
| article_number | 3181 | 665 | 678 | 3.9 |
| arrangement | 644 | 122 | 153 | 0.8 |
| Total | 80831 | 16698 | 17894 | 100 |
Data encoding
Transcriptions with entities are encoded in the labels.json JSON file. Special tokens are used to represent named entities. Please not that there are only opening NER tokens: each entity spans all words until the next entity starts.
| Entities | Special token | Symbol unicode |
|---|---|---|
| date | ⓓ | \u24d3 |
| title | ⓘ | \u24d8 |
| serie | ⓢ | \u24e2 |
| analysis | ⓒ | \u24d2 |
| volume_number | ⓟ | \u24df |
| article_number | ⓐ | \u24d0 |
| arrangement | ⓥ | \u24e5 |
Cite us!
The dataset is presented in details in the following article:
@article{simara2023,
author = {Solène Tarride and Mélodie Boillet and Jean-François Moufflet and Christopher Kermorvant},
title = {SIMARA: a database for key-value information extraction from full-page handwritten documents},
year = {2023},
journal={Proceedings of the 17th International Conference on Document Analysis and Recognition},
}
Benchmarks archive 2025-07-28
All 2 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.
| First row (archive order) | Paper | Code | ||||
|---|---|---|---|---|---|---|
| Handwritten Text Recognition | SIMARA | DAN CER (%) 6.46 | SIMARA: a database for key-value information extraction... | — | 1 | Compare |
| Key Information Extraction | SIMARA | DAN F1 (%) 95.05 | SIMARA: a database for key-value information extraction... | — | 1 | Compare |
Papers archive 2025-07-28
1 shown of 1 paper with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 2. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.
| Date | Samples run Syntology | |||
|---|---|---|---|---|
| SIMARA: a database for key-value information extraction from full pages | 0 | 2 | 26 Apr 2023 | not harvested |
Dataset loaders archive 2025-07-28
No loader listed in the archive.
Tasks archive 2025-07-28
License archive 2025-07-28
Creative Commons Attribution 4.0 International
Modalities archive 2025-07-28
Languages archive 2025-07-28
Variants archive 2025-07-28
- SIMARA
1 variant name, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections