{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/vision-grid-transformer-for-document-layout","title":"Vision Grid Transformer for Document Layout Analysis","arxiv_id":"2308.14978","date":"2023-08-29","proceeding":"ICCV 2023 1","authors":["Cheng Da","Chuwei Luo","Qi Zheng","Cong Yao"],"abstract":"Document pre-trained models and grid-based models have proven to be very effective on various tasks in Document AI. However, for the document layout analysis (DLA) task, existing document pre-trained models, even those pre-trained in a multi-modal fashion, usually rely on either textual features or visual features. Grid-based models for DLA are multi-modality but largely neglect the effect of pre-training. To fully leverage multi-modal information and exploit pre-training techniques to learn better representation for DLA, in this paper, we present VGT, a two-stream Vision Grid Transformer, in which Grid Transformer (GiT) is proposed and pre-trained for 2D token-level and segment-level semantic understanding. Furthermore, a new dataset named D$^4$LA, which is so far the most diverse and detailed manually-annotated benchmark for document layout analysis, is curated and released. Experiment results have illustrated that the proposed VGT model achieves new state-of-the-art results on DLA tasks, e.g. PubLayNet ($95.7\\%$$\\rightarrow$$96.2\\%$), DocBank ($79.6\\%$$\\rightarrow$$84.1\\%$), and D$^4$LA ($67.7\\%$$\\rightarrow$$68.8\\%$). The code and models as well as the D$^4$LA dataset will be made publicly available ~\\url{https://github.com/AlibabaResearch/AdvancedLiterateMachinery}.","url_abs":"https://arxiv.org/abs/2308.14978v1","url_pdf":"https://arxiv.org/pdf/2308.14978v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"vision-grid-transformer-for-document-layout","repo_url":"https://github.com/alibabaresearch/advancedliteratemachinery","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"document-ai","task_name":"Document AI"},{"task_slug":"document-layout-analysis","task_name":"Document Layout Analysis"},{"task_slug":"optical-character-recognition","task_name":"Optical Character Recognition (OCR)"},{"task_slug":"document-understanding","task_name":"document understanding"}],"methods":[{"method_slug":"absolute-position-encodings","method_name":"Absolute Position Encodings"},{"method_slug":"adam","method_name":"Adam"},{"method_slug":"attention","method_name":"Attention"},{"method_slug":"bpe","method_name":"BPE"},{"method_slug":"dla","method_name":"DLA"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"dropout","method_name":"Dropout"},{"method_slug":"label-smoothing","method_name":"Label Smoothing"},{"method_slug":"layer-normalization","method_name":"Layer Normalization"},{"method_slug":"linear-layer","method_name":"Linear Layer"},{"method_slug":"multi-head-attention","method_name":"Multi-Head Attention"},{"method_slug":"position-wise-feed-forward-layer","method_name":"Position-Wise Feed-Forward Layer"},{"method_slug":"residual-connection","method_name":"Residual Connection"},{"method_slug":"softmax","method_name":"Softmax"},{"method_slug":"transformer","method_name":"Transformer"}],"datasets_introduced":[{"slug":"d4la","name":"D4LA","full_name":""}],"methods_introduced":[],"results":[{"leaderboard":"/sota/document-layout-analysis-on-d4la","task":"Document Layout Analysis","dataset":"D4LA","model":"VGT","rank_in_archive_order":3,"of":3,"metrics":{" mAP":"68.8","Model Parameters":"174M"},"uses_additional_data":false},{"leaderboard":"/sota/document-layout-analysis-on-publaynet-val","task":"Document Layout Analysis","dataset":"PubLayNet val","model":"VGT","rank_in_archive_order":1,"of":15,"metrics":{"Figure":"0.971","List":"0.968","Overall":"0.962","Table":"0.981","Text":"0.950","Title":"0.939"},"uses_additional_data":false},{"leaderboard":"/sota/document-layout-analysis-on-publaynet-val","task":"Document Layout Analysis","dataset":"PubLayNet val","model":"ResNext-101-32×8d","rank_in_archive_order":9,"of":15,"metrics":{"Figure":"0.968","List":"0.940","Overall":"0.935","Table":"0.976","Text":"0.930","Title":"0.862"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=2308.14978","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}