Papers › Vision Grid Transformer for Document Layout Analysis

Vision Grid Transformer for Document Layout Analysis

29 Aug 2023ICCV 2023 1arXiv:2308.14978archive 2025-07-28

Cheng Da, Chuwei Luo, Qi Zheng, Cong Yao

Document pre-trained models and grid-based models have proven to be very effective on various tasks in Document AI. However, for the document layout analysis (DLA) task, existing document pre-trained models, even those pre-trained in a multi-modal fashion, usually rely on either textual features or visual features. Grid-based models for DLA are multi-modality but largely neglect the effect of pre-training. To fully leverage multi-modal information and exploit pre-training techniques to learn better representation for DLA, in this paper, we present VGT, a two-stream Vision Grid Transformer, in which Grid Transformer (GiT) is proposed and pre-trained for 2D token-level and segment-level semantic understanding. Furthermore, a new dataset named D⁴LA, which is so far the most diverse and detailed manually-annotated benchmark for document layout analysis, is curated and released. Experiment results have illustrated that the proposed VGT model achieves new state-of-the-art results on DLA tasks, e.g. PubLayNet ($95.7\%→96.2\%$), DocBank ($79.6\%→84.1\%), and D^4$LA ($67.7\%→68.8\%). The code and models as well as the D^4$LA dataset will be made publicly available ~\url{https://github.com/AlibabaResearch/AdvancedLiterateMachinery}.

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

alibabaresearch/advancedliteratemachinery officialmentioned in papermentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Document AIDocument Layout AnalysisOptical Character Recognition (OCR)document understanding

Datasets

Introduced by this paper, per the archive.

D4LA

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Document Layout Analysis D4LA VGT mAP 68.8 #3 of 3 Archive leaderboard report
Document Layout Analysis D4LA VGT Model Parameters 174M #3 of 3 Archive leaderboard report
Document Layout Analysis PubLayNet val VGT Figure 0.971 #1 of 15 Archive leaderboard report
Document Layout Analysis PubLayNet val VGT List 0.968 #1 of 15 Archive leaderboard report
Document Layout Analysis PubLayNet val VGT Overall 0.962 #1 of 15 Archive leaderboard report
Document Layout Analysis PubLayNet val VGT Table 0.981 #1 of 15 Archive leaderboard report
Document Layout Analysis PubLayNet val VGT Text 0.950 #1 of 15 Archive leaderboard report
Document Layout Analysis PubLayNet val VGT Title 0.939 #1 of 15 Archive leaderboard report
Document Layout Analysis PubLayNet val ResNext-101-32×8d Figure 0.968 #9 of 15 Archive leaderboard report
Document Layout Analysis PubLayNet val ResNext-101-32×8d List 0.940 #9 of 15 Archive leaderboard report
Document Layout Analysis PubLayNet val ResNext-101-32×8d Overall 0.935 #9 of 15 Archive leaderboard report
Document Layout Analysis PubLayNet val ResNext-101-32×8d Table 0.976 #9 of 15 Archive leaderboard report
Document Layout Analysis PubLayNet val ResNext-101-32×8d Text 0.930 #9 of 15 Archive leaderboard report
Document Layout Analysis PubLayNet val ResNext-101-32×8d Title 0.862 #9 of 15 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Absolute Position EncodingsAdamAttentionBPEDLADense ConnectionsDropoutLabel SmoothingLayer NormalizationLinear LayerMulti-Head AttentionPosition-Wise Feed-Forward LayerResidual ConnectionSoftmaxTransformer

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections