Papers › CorDeep and the Sacrobosco Dataset: Detection of Visual Elements in Historical Documents

CorDeep and the Sacrobosco Dataset: Detection of Visual Elements in Historical Documents

15 Oct 2022Journal of Imaging 2022 10archive 2025-07-28

Jochen Büttner, Julius Martinetz, Hassan El-Hajj, Matteo Valleriani

Recent advances in object detection facilitated by deep learning have led to numerous solutions in a myriad of fields ranging from medical diagnosis to autonomous driving. However, historical research is yet to reap the benefits of such advances. This is generally due to the low number of large, coherent, and annotated datasets of historical documents, as well as the overwhelming focus on Optical Character Recognition to support the analysis of historical documents. In this paper, we highlight the importance of visual elements, in particular illustrations in historical documents, and offer a public multi-class historical visual element dataset based on the Sphaera corpus. Additionally, we train an image extraction model based on YOLO architecture and publish it through a publicly available web-service to detect and extract multi-class images from historical documents in an effort to bridge the gap between traditional and computational approaches in historical studies.

PaperPDF

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Object DetectionOptical Character RecognitionSemantic Segmentationobject-detection

Datasets

Introduced by this paper, per the archive.

S-VED

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections