{"url":"/dataset/belfort","name":"Belfort","full_name":"The Belfort dataset: Handwritten Text Recognition from Crowdsourced Annotations","description_markdown":"# The Belfort dataset\r\n\r\nThis dataset includes minutes of Belfort municipal council drawn up between 1790 and 1946. Documents include deliberations, lists of councillors, convocations, and agendas. It includes 24,105 text-line images that were automatically detected from pages. Up to 4 transcriptions are available for each line image: two from humans, and two from automatic models.\r\n\r\nFiles are organized in three folders: `Images`, `Transcriptions`, and `Partitions`.\r\n\r\n## Images\r\n\r\nThe dataset include 24,105 text-line images that were automatically detected using a generic [Doc-UFCN](https://pypi.org/project/doc-ufcn/) model, and resized to a fixed height of 128 pixels.\r\n\r\n## Transcriptions\r\nUp to 4 transcriptions are available for each image, as summarized in the following table:\r\n\r\n|   Folder   \t| N transcriptions \t| Description                 | Comments                                                                          |\r\n|:----------:\t|-----------------:\t|-----------------------------|-----------------------------------------------------------------------------------|\r\n| callico_1/ \t|           24,105 \t| Human annotation n°1        | All lines have at least one human annotation                                      |\r\n| callico_2/ \t|            8,878 \t| Human annotation n°2        | Only 33% of lines have two different human annotations                            |\r\n| dan/       \t|           24,102 \t| DAN automatic model         | 3 images have empty transcriptions (no text was predicted by the model)           |\r\n| pylaia/    \t|           23,536 \t| PyLaia automatic model      | 569 images have empty transcriptions (no text was predicted by the model)         |\r\n| rasa/     \t|           23,287 \t| RASA aggregation algorithm  | 818 images have empty transcriptions                                              |\r\n| rover/    \t|           24,104 \t| ROVER aggregation algorithm | 1 image has an empty transcription                                                |\r\n\r\n## Data partition\r\n\r\nWe provide two distinct splits, both of them containing 19,013 training images, 2,262 validation images and 2,830 test images.\r\n\r\n* The *Agreement-based split* ensures the reliability of the test set:\r\n    * The test set includes lines with perfect agreement between human annotators (Character Error Rate = 0%);\r\n    * The validation set includes lines with good agreement between human annotators (0% < Character Error Rate < 5%);\r\n    * The training set includes all the other lines.\r\n* The *Random split* is randomized.\r\n\r\n## Evaluation\r\n\r\nEvaluation results in the paper are computed by comparing predictions to human annotations.\r\nAutomatic and aggregated transcriptions are only used during model training.","description_withheld":null,"homepage":"https://zenodo.org/record/8041668","introduced_date":"2023-06-15","introduced_date_note":null,"introduced_by":null,"license":{"name":"Creative Commons Attribution 4.0 International","url":"https://creativecommons.org/licenses/by/4.0/legalcode"},"modalities":[{"name":"Images","url":"/datasets/modality/images"},{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Handwritten Text Recognition","url":"/task/handwritten-text-recognition","datasets_with_task":"/datasets/task/handwritten-text-recognition"}],"languages":[{"name":"French","url":"/datasets/language/french"}],"variants":["Belfort"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/handwritten-text-recognition-on-belfort","task":"Handwritten Text Recognition","dataset_variant":"Belfort","rows":4,"metrics":["CER (%)","WER (%)"],"first_row_in_archive_order":{"model":"PyLaia (all transcriptions + agreement-based split)","paper":"/paper/handwritten-text-recognition-from-1","metrics":{"CER (%)":"4.34","WER (%)":"15.14"},"code_links":[]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/handwritten-text-recognition-from-1","title":"Handwritten Text Recognition from Crowdsourced Annotations","date":"2023-06-19","rows_on_this_dataset":4,"code_links":0,"syntology":null}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}