{"url":"/dataset/lam-line-level","name":"LAM(line-level)","full_name":"The LAM Dataset: A Novel Benchmark for Line-Level Handwritten Text Recognition","description_markdown":"Handwritten Text Recognition (HTR) is an open\r\nproblem at the intersection of Computer Vision and Natural\r\nLanguage Processing. The main challenges, when dealing with\r\nhistorical manuscripts, are due to the preservation of the paper\r\nsupport, the variability of the handwriting – even of the same\r\nauthor over a wide time-span – and the scarcity of data from\r\nancient, poorly represented languages. With the aim of fostering\r\nthe research on this topic, in this paper we present the Ludovico\r\nAntonio Muratori (LAM) dataset, a large line-level HTR dataset\r\nof Italian ancient manuscripts edited by a single author over\r\n60 years. The dataset comes in two configurations: a basic\r\nsplitting and a date-based splitting which takes into account the\r\nage of the author. The first setting is intended to study HTR\r\non ancient documents in Italian, while the second focuses on\r\nthe ability of HTR systems to recognize text written by the\r\nsame writer in time periods for which training data are not\r\navailable. For both configurations, we analyze quantitative and\r\nqualitative characteristics, also with respect to other line-level\r\nHTR benchmarks, and present the recognition performance of\r\nstate-of-the-art HTR architectures. The dataset is available for\r\ndownload at https://aimagelab.ing.unimore.it/go/lam.","description_withheld":null,"homepage":"https://arxiv.org/pdf/2208.07682","introduced_date":"2022-08-16","introduced_date_note":null,"introduced_by":null,"license":null,"modalities":[{"name":"Images","url":"/datasets/modality/images"},{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Handwritten Text Recognition","url":"/task/handwritten-text-recognition","datasets_with_task":"/datasets/task/handwritten-text-recognition"}],"languages":[{"name":"Italian","url":"/datasets/language/italian"}],"variants":["LAM(line-level)"],"data_loaders":[],"num_papers_in_archive":4,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/handwritten-text-recognition-on-lam-line","task":"Handwritten Text Recognition","dataset_variant":"LAM(line-level)","rows":6,"metrics":["Test CER","Test WER"],"first_row_in_archive_order":{"model":"HTR-VT","paper":"/paper/htr-vt-handwritten-text-recognition-with","metrics":{"Test CER":"2.8","Test WER":"7.4"},"code_links":[{"title":"yutingli0606/htr-vt","url":"https://github.com/yutingli0606/htr-vt"},{"title":"mortezagolzan/Hand_Writing_Text_Extarction","url":"https://github.com/mortezagolzan/Hand_Writing_Text_Extarction"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/htr-vt-handwritten-text-recognition-with","title":"HTR-VT: Handwritten Text Recognition with Vision Transformer","date":"2024-09-13","rows_on_this_dataset":1,"code_links":2,"syntology":null},{"paper":"/paper/trocr-transformer-based-optical-character","title":"TrOCR: Transformer-based Optical Character Recognition with Pre-trained Models","date":"2021-09-21","rows_on_this_dataset":1,"code_links":8,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":6,"samples_ran":0,"samples_unverified":6,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/recurrence-free-unconstrained-handwritten","title":"Recurrence-free unconstrained handwritten text recognition using gated fully convolutional network","date":"2020-12-09","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/origaminet-weakly-supervised-segmentation-1","title":"OrigamiNet: Weakly-Supervised, Segmentation-Free, One-Step, Full Page Text Recognition by learning to unfold","date":"2020-06-12","rows_on_this_dataset":3,"code_links":8,"syntology":null}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":1,"samples_harvested":6,"samples_ran":0,"samples_unverified":6,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":1,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}