{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/analysis-of-convolutional-neural-networks-for","title":"Analysis of Convolutional Neural Networks for Document Image Classification","arxiv_id":"1708.03273","date":"2017-08-10","proceeding":null,"authors":["Chris Tensmeyer","Tony Martinez"],"abstract":"Convolutional Neural Networks (CNNs) are state-of-the-art models for document\nimage classification tasks. However, many of these approaches rely on\nparameters and architectures designed for classifying natural images, which\ndiffer from document images. We question whether this is appropriate and\nconduct a large empirical study to find what aspects of CNNs most affect\nperformance on document images. Among other results, we exceed the\nstate-of-the-art on the RVL-CDIP dataset by using shear transform data\naugmentation and an architecture designed for a larger input image.\nAdditionally, we analyze the learned features and find evidence that CNNs\ntrained on RVL-CDIP learn region-specific layout features.","url_abs":"http://arxiv.org/abs/1708.03273v1","url_pdf":"http://arxiv.org/pdf/1708.03273v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"classification-1","task_name":"Classification"},{"task_slug":"data-augmentation","task_name":"Data Augmentation"},{"task_slug":"document-image-classification","task_name":"Document Image Classification"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"image-classification","task_name":"Image Classification"},{"task_slug":"document-image-classification","task_name":"document-image-classification"},{"task_slug":"image-classification","task_name":"image-classification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/document-image-classification-on-rvl-cdip","task":"Document Image Classification","dataset":"RVL-CDIP","model":"AlexNet + spatial pyramidal pooling + image resizing","rank_in_archive_order":29,"of":31,"metrics":{"Accuracy":"90.94%"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/1708.03273","atlas_url":"https://app.syntology.ai/?focus=1708.03273","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}