{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/structextv2-masked-visual-textual-prediction","title":"StrucTexTv2: Masked Visual-Textual Prediction for Document Image Pre-training","arxiv_id":"2303.00289","date":"2023-03-01","proceeding":null,"authors":["Yuechen Yu","Yulin Li","Chengquan Zhang","Xiaoqiang Zhang","Zengyuan Guo","Xiameng Qin","Kun Yao","Junyu Han","Errui Ding","Jingdong Wang"],"abstract":"In this paper, we present StrucTexTv2, an effective document image pre-training framework, by performing masked visual-textual prediction. It consists of two self-supervised pre-training tasks: masked image modeling and masked language modeling, based on text region-level image masking. The proposed method randomly masks some image regions according to the bounding box coordinates of text words. The objectives of our pre-training tasks are reconstructing the pixels of masked image regions and the corresponding masked tokens simultaneously. Hence the pre-trained encoder can capture more textual semantics in comparison to the masked image modeling that usually predicts the masked image patches. Compared to the masked multi-modal modeling methods for document image understanding that rely on both the image and text modalities, StrucTexTv2 models image-only input and potentially deals with more application scenarios free from OCR pre-processing. Extensive experiments on mainstream benchmarks of document image understanding demonstrate the effectiveness of StrucTexTv2. It achieves competitive or even new state-of-the-art performance in various downstream tasks such as image classification, layout analysis, table structure recognition, document OCR, and information extraction under the end-to-end scenario.","url_abs":"https://arxiv.org/abs/2303.00289v1","url_pdf":"https://arxiv.org/pdf/2303.00289v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"structextv2-masked-visual-textual-prediction","repo_url":"https://github.com/PaddlePaddle/VIMER/tree/main/StrucTexT/v2","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"paddle","reach":null}],"tasks":[{"task_slug":"document-image-classification","task_name":"Document Image Classification"},{"task_slug":"image-classification","task_name":"Image Classification"},{"task_slug":"language-modeling","task_name":"Language Modeling"},{"task_slug":"language-modelling","task_name":"Language Modelling"},{"task_slug":"masked-language-modeling","task_name":"Masked Language Modeling"},{"task_slug":"optical-character-recognition","task_name":"Optical Character Recognition (OCR)"},{"task_slug":"semantic-entity-labeling","task_name":"Semantic entity labeling"},{"task_slug":"image-classification","task_name":"image-classification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/document-image-classification-on-rvl-cdip","task":"Document Image Classification","dataset":"RVL-CDIP","model":"StrucTexTv2 (large)","rank_in_archive_order":14,"of":31,"metrics":{"Accuracy":"94.62%","Parameters":"238M"},"uses_additional_data":false},{"leaderboard":"/sota/document-image-classification-on-rvl-cdip","task":"Document Image Classification","dataset":"RVL-CDIP","model":"StrucTexTv2 (small)","rank_in_archive_order":18,"of":31,"metrics":{"Accuracy":"93.4%","Parameters":"28M"},"uses_additional_data":false},{"leaderboard":"/sota/semantic-entity-labeling-on-funsd","task":"Semantic entity labeling","dataset":"FUNSD","model":"StrucTexTv2 (large)","rank_in_archive_order":7,"of":15,"metrics":{"F1":"91.82"},"uses_additional_data":false},{"leaderboard":"/sota/semantic-entity-labeling-on-funsd","task":"Semantic entity labeling","dataset":"FUNSD","model":"StrucTexTv2 (small)","rank_in_archive_order":9,"of":15,"metrics":{"F1":"89.23"},"uses_additional_data":false},{"leaderboard":"/sota/table-recognition-on-wtw","task":"Table Recognition","dataset":"WTW","model":"StrucTexTv2 (small)","rank_in_archive_order":1,"of":1,"metrics":{"F1":"78.9%"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2303.00289","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}