{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/read-bad-a-new-dataset-and-evaluation-scheme","title":"READ-BAD: A New Dataset and Evaluation Scheme for Baseline Detection in Archival Documents","arxiv_id":"1705.03311","date":"2017-05-09","proceeding":null,"authors":["Tobias Grüning","Roger Labahn","Markus Diem","Florian Kleber","Stefan Fiel"],"abstract":"Text line detection is crucial for any application associated with Automatic\nText Recognition or Keyword Spotting. Modern algorithms perform good on\nwell-established datasets since they either comprise clean data or\nsimple/homogeneous page layouts. We have collected and annotated 2036 archival\ndocument images from different locations and time periods. The dataset contains\nvarying page layouts and degradations that challenge text line segmentation\nmethods. Well established text line segmentation evaluation schemes such as the\nDetection Rate or Recognition Accuracy demand for binarized data that is\nannotated on a pixel level. Producing ground truth by these means is laborious\nand not needed to determine a method's quality. In this paper we propose a new\nevaluation scheme that is based on baselines. The proposed scheme has no need\nfor binarization and it can handle skewed as well as rotated text lines. The\nICDAR 2017 Competition on Baseline Detection and the ICDAR 2017 Competition on\nLayout Analysis for Challenging Medieval Manuscripts used this evaluation\nscheme. Finally, we present results achieved by a recently published text line\ndetection algorithm.","url_abs":"http://arxiv.org/abs/1705.03311v2","url_pdf":"http://arxiv.org/pdf/1705.03311v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"read-bad-a-new-dataset-and-evaluation-scheme","repo_url":"https://github.com/Transkribus/TranskribusBaseLineEvaluationScheme","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null},{"paper_slug":"read-bad-a-new-dataset-and-evaluation-scheme","repo_url":"https://github.com/TobiasGruening/ARU-Net","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null},{"paper_slug":"read-bad-a-new-dataset-and-evaluation-scheme","repo_url":"https://github.com/dhlab-epfl/dhSegment","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null},{"paper_slug":"read-bad-a-new-dataset-and-evaluation-scheme","repo_url":"https://github.com/imagine5am/ARU-Net","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"binarization","task_name":"Binarization"},{"task_slug":"keyword-spotting","task_name":"Keyword Spotting"},{"task_slug":"line-detection","task_name":"Line Detection"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}