{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/cutting-the-error-by-half-investigation-of","title":"Cutting the Error by Half: Investigation of Very Deep CNN and Advanced Training Strategies for Document Image Classification","arxiv_id":"1704.03557","date":"2017-04-11","proceeding":null,"authors":["Muhammad Zeshan Afzal","Andreas Kölsch","Sheraz Ahmed","Marcus Liwicki"],"abstract":"We present an exhaustive investigation of recent Deep Learning architectures,\nalgorithms, and strategies for the task of document image classification to\nfinally reduce the error by more than half. Existing approaches, such as the\nDeepDocClassifier, apply standard Convolutional Network architectures with\ntransfer learning from the object recognition domain. The contribution of the\npaper is threefold: First, it investigates recently introduced very deep neural\nnetwork architectures (GoogLeNet, VGG, ResNet) using transfer learning (from\nreal images). Second, it proposes transfer learning from a huge set of document\nimages, i.e. 400,000 documents. Third, it analyzes the impact of the amount of\ntraining data (document images) and other parameters to the classification\nabilities. We use two datasets, the Tobacco-3482 and the large-scale RVL-CDIP\ndataset. We achieve an accuracy of 91.13% for the Tobacco-3482 dataset while\nearlier approaches reach only 77.6%. Thus, a relative error reduction of more\nthan 60% is achieved. For the large dataset RVL-CDIP, an accuracy of 90.97% is\nachieved, corresponding to a relative error reduction of 11.5%.","url_abs":"http://arxiv.org/abs/1704.03557v1","url_pdf":"http://arxiv.org/pdf/1704.03557v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"cutting-the-error-by-half-investigation-of","repo_url":"https://github.com/BordiaS/layoutlm","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"cutting-the-error-by-half-investigation-of","repo_url":"https://github.com/iamarjunchandra/LayoutLM-Form-Understanding---Sequence-Labeling","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"cutting-the-error-by-half-investigation-of","repo_url":"https://github.com/microsoft/unilm/tree/master/layoutlm","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"cutting-the-error-by-half-investigation-of","repo_url":"https://github.com/tuannamnguyen93/DFKI_test_PhD","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"unanswered"}},{"paper_slug":"cutting-the-error-by-half-investigation-of","repo_url":"https://github.com/tuannamnguyen93/test_PhD","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"document-image-classification","task_name":"Document Image Classification"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"image-classification","task_name":"Image Classification"},{"task_slug":"object-recognition","task_name":"Object Recognition"},{"task_slug":"transfer-learning","task_name":"Transfer Learning"},{"task_slug":"document-image-classification","task_name":"document-image-classification"},{"task_slug":"image-classification","task_name":"image-classification"}],"methods":[{"method_slug":"1x1-convolution","method_name":"1x1 Convolution"},{"method_slug":"average-pooling","method_name":"Average Pooling"},{"method_slug":"batch-normalization","method_name":"Batch Normalization"},{"method_slug":"bottleneck-residual-block","method_name":"Bottleneck Residual Block"},{"method_slug":"convolution","method_name":"Convolution"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"dropout","method_name":"Dropout"},{"method_slug":"global-average-pooling","method_name":"Global Average Pooling"},{"method_slug":"kaiming-initialization","method_name":"Kaiming Initialization"},{"method_slug":"max-pooling","method_name":"Max Pooling"},{"method_slug":"relu","method_name":"ReLU"},{"method_slug":"residual-block","method_name":"Residual Block"},{"method_slug":"residual-connection","method_name":"Residual Connection"},{"method_slug":"softmax","method_name":"Softmax"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/document-image-classification-on-rvl-cdip","task":"Document Image Classification","dataset":"RVL-CDIP","model":"Transfer Learning from AlexNet, VGG-16, GoogLeNet and ResNet50","rank_in_archive_order":28,"of":31,"metrics":{"Accuracy":"90.97%"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1704.03557","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}