{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/docxclassifier-high-performance-explainable","title":"DocXClassifier: High Performance Explainable Deep Network for Document Image Classification","arxiv_id":null,"date":"2022-03-17","proceeding":"TechArXiv 2022 3","authors":["Saifullah","Stefan Agne","Andreas Dengel","Sheraz Ahmed"],"abstract":"Convolutional Neural Networks (ConvNets) have been thoroughly researched for document\r\nimage classification and are known for their exceptional performance in unimodal image-based document\r\nclassification. Recently, however, there has been a sudden shift in the field towards multimodal approaches\r\nthat simultaneously learn from the visual and textual features of the documents. While this has led to\r\nsignificant advances in the field, it has also led to a waning interest in improving pure ConvNets-based\r\napproaches. This is not desirable, as many of the multimodal approaches still use ConvNets as their visual\r\nbackbone, and thus improving ConvNets is essential to improving these approaches. In this paper, we present\r\nDocXClassifier, a ConvNet-based approach that, using state-of-the-art model design patterns together with\r\nmodern data augmentation and training strategies, not only achieves significant performance improvements\r\nin image-based document classification, but also outperforms some of the recently proposed multimodal\r\napproaches. Moreover, DocXClassifier is capable of generating transformer-like attention maps, which\r\nmakes it inherently interpretable, a property not found in previous image-based classification models. Our\r\napproach achieves a new peak performance in image-based classification on two popular document datasets,\r\nnamely RVL-CDIP and Tobacco3482, with a top-1 classification accuracy of 94.17% and 95.57% on the two\r\ndatasets, respectively. Moreover, it sets a new record for the highest image-based classification accuracy of\r\n90.14% on Tobacco3482 without transfer learning from RVL-CDIP. Finally, our proposed model may serve\r\nas a powerful visual backbone for future multimodal approaches, by providing much richer visual features\r\nthan existing counterparts.","url_abs":"https://www.techrxiv.org/articles/preprint/DocXClassifier_High_Performance_Explainable_Deep_Network_for_Document_Image_Classification/19310489","url_pdf":"https://www.techrxiv.org/articles/preprint/DocXClassifier_High_Performance_Explainable_Deep_Network_for_Document_Image_Classification/19310489","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"docxclassifier-high-performance-explainable","repo_url":"https://github.com/saifullah3396/docxclassifier","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"classification-1","task_name":"Classification"},{"task_slug":"data-augmentation","task_name":"Data Augmentation"},{"task_slug":"document-classification","task_name":"Document Classification"},{"task_slug":"document-image-classification","task_name":"Document Image Classification"},{"task_slug":"image-classification","task_name":"Image Classification"},{"task_slug":"transfer-learning","task_name":"Transfer Learning"},{"task_slug":"high","task_name":"Vocal Bursts Intensity Prediction"},{"task_slug":"document-image-classification","task_name":"document-image-classification"},{"task_slug":"image-classification","task_name":"image-classification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/document-image-classification-on-rvl-cdip","task":"Document Image Classification","dataset":"RVL-CDIP","model":"DocXClassifier-B","rank_in_archive_order":17,"of":31,"metrics":{"Accuracy":"94.00%","Parameters":"95.4M"},"uses_additional_data":false},{"leaderboard":"/sota/document-image-classification-on-tobacco-3482","task":"Document Image Classification","dataset":"Tobacco-3482","model":"DocXClassifier-L","rank_in_archive_order":1,"of":10,"metrics":{"Accuracy":"95.57"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}