{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-comprehensive-study-of-imagenet-pre","title":"A Comprehensive Study of ImageNet Pre-Training for Historical Document Image Analysis","arxiv_id":"1905.09113","date":"2019-05-22","proceeding":null,"authors":["Linda Studer","Michele Alberti","Vinaychandran Pondenkandath","Pinar Goktepe","Thomas Kolonko","Andreas Fischer","Marcus Liwicki","Rolf Ingold"],"abstract":"Automatic analysis of scanned historical documents comprises a wide range of image analysis tasks, which are often challenging for machine learning due to a lack of human-annotated learning samples. With the advent of deep neural networks, a promising way to cope with the lack of training data is to pre-train models on images from a different domain and then fine-tune them on historical documents. In the current research, a typical example of such cross-domain transfer learning is the use of neural networks that have been pre-trained on the ImageNet database for object recognition. It remains a mostly open question whether or not this pre-training helps to analyse historical documents, which have fundamentally different image properties when compared with ImageNet. In this paper, we present a comprehensive empirical survey on the effect of ImageNet pre-training for diverse historical document analysis tasks, including character recognition, style classification, manuscript dating, semantic segmentation, and content-based retrieval. While we obtain mixed results for semantic segmentation at pixel-level, we observe a clear trend across different network architectures that ImageNet pre-training has a positive effect on classification as well as content-based retrieval.","url_abs":"https://arxiv.org/abs/1905.09113v1","url_pdf":"https://arxiv.org/pdf/1905.09113v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"image-classification","task_name":"Image Classification"},{"task_slug":"object-recognition","task_name":"Object Recognition"},{"task_slug":"open-question","task_name":"Open-Ended Question Answering"},{"task_slug":"retrieval","task_name":"Retrieval"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"},{"task_slug":"transfer-learning","task_name":"Transfer Learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/image-classification-on-kuzushiji-mnist","task":"Image Classification","dataset":"Kuzushiji-MNIST","model":"Resnet-152","rank_in_archive_order":9,"of":26,"metrics":{"Accuracy":"98.79"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/1905.09113","atlas_url":"https://app.syntology.ai/?focus=1905.09113","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}