Papers › LayoutMask: Enhance Text-Layout Interaction in Multi-modal Pre-training for Document...

LayoutMask: Enhance Text-Layout Interaction in Multi-modal Pre-training for Document Understanding

30 May 2023arXiv:2305.18721archive 2025-07-28

Yi Tu, Ya Guo, Huan Chen, Jinyang Tang

Visually-rich Document Understanding (VrDU) has attracted much research attention over the past years. Pre-trained models on a large number of document images with transformer-based backbones have led to significant performance gains in this field. The major challenge is how to fusion the different modalities (text, layout, and image) of the documents in a unified model with different pre-training tasks. This paper focuses on improving text-layout interactions and proposes a novel multi-modal pre-training model, LayoutMask. LayoutMask uses local 1D position, instead of global 1D position, as layout input and has two pre-training objectives: (1) Masked Language Modeling: predicting masked tokens with two novel masking strategies; (2) Masked Position Modeling: predicting masked 2D positions to improve layout representation learning. LayoutMask can enhance the interactions between text and layout modalities in a unified model and produce adaptive and robust multi-modal representations for downstream tasks. Experimental results show that our proposed method can achieve state-of-the-art results on a wide variety of VrDU problems, including form understanding, receipt understanding, and document image classification.

PaperPDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Document Image ClassificationImage ClassificationKey Information ExtractionLanguage ModelingLanguage ModellingMasked Language ModelingNamed Entity Recognition (NER)Representation LearningSemantic entity labelingdocument understandingdocument-image-classificationimage-classification

1 archive task tag without a task page not shown.

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Key Information Extraction CORD LayoutMask (large) F1 97.19 #4 of 9 Archive leaderboard report
Key Information Extraction CORD LayoutMask (base) F1 96.99 #5 of 9 Archive leaderboard report
Named Entity Recognition (NER) CORD-r LayoutMask F1 81.84 #4 of 4 Archive leaderboard report
Named Entity Recognition (NER) FUNSD-r LayoutMask F1 77.10 #4 of 4 Archive leaderboard report
Semantic entity labeling FUNSD LayoutMask (large) F1 93.20 #1 of 15 Archive leaderboard report
Semantic entity labeling FUNSD LayoutMask (base) F1 92.91 #3 of 15 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections