Papers › LayoutLMv3: Pre-training for Document AI with Unified Text and Image Masking

LayoutLMv3: Pre-training for Document AI with Unified Text and Image Masking

18 Apr 2022arXiv:2204.08387archive 2025-07-28

Yupan Huang, Tengchao Lv, Lei Cui, Yutong Lu, Furu Wei

Self-supervised pre-training techniques have achieved remarkable progress in Document AI. Most multimodal pre-trained models use a masked language modeling objective to learn bidirectional representations on the text modality, but they differ in pre-training objectives for the image modality. This discrepancy adds difficulty to multimodal representation learning. In this paper, we propose \textbf{LayoutLMv3} to pre-train multimodal Transformers for Document AI with unified text and image masking. Additionally, LayoutLMv3 is pre-trained with a word-patch alignment objective to learn cross-modal alignment by predicting whether the corresponding image patch of a text word is masked. The simple unified architecture and training objectives make LayoutLMv3 a general-purpose pre-trained model for both text-centric and image-centric Document AI tasks. Experimental results show that LayoutLMv3 achieves state-of-the-art performance not only in text-centric tasks, including form understanding, receipt understanding, and document visual question answering, but also in image-centric tasks such as document image classification and document layout analysis. The code and models are publicly available at \url{https://aka.ms/layoutlmv3}.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

microsoft/unilm officialpytorch report
huggingface/transformers mentioned on GitHubpytorch report
pwc-1/Paper-9 mindspore report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Document AIDocument Image ClassificationDocument Layout AnalysisEntity LinkingImage ClassificationKey Information ExtractionKey-value Pair ExtractionLanguage ModelingLanguage ModellingMasked Language ModelingNamed Entity Recognition (NER)Question AnsweringRelation ExtractionRepresentation LearningSemantic entity labelingVisual Question AnsweringVisual Question Answering (VQA)cross-modal alignmentdocument-image-classificationimage-classification

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Document AI EPHOIE LayoutLMv3 Average F1 99.21 #1 of 1 Archive leaderboard report
Document Image Classification RVL-CDIP LayoutLMV3Large Accuracy 95.93% #4 of 31 Archive leaderboard report
Document Image Classification RVL-CDIP LayoutLMV3Large Parameters 368M #4 of 31 Archive leaderboard report
Document Image Classification RVL-CDIP LayoutLMv3BASE Accuracy 95.44% #9 of 31 Archive leaderboard report
Document Image Classification RVL-CDIP LayoutLMv3BASE Parameters 133M #9 of 31 Archive leaderboard report
Document Layout Analysis PubLayNet val LayoutLMv3-B Figure 0.970 #5 of 15 Archive leaderboard report
Document Layout Analysis PubLayNet val LayoutLMv3-B List 0.955 #5 of 15 Archive leaderboard report
Document Layout Analysis PubLayNet val LayoutLMv3-B Overall 0.951 #5 of 15 Archive leaderboard report
Document Layout Analysis PubLayNet val LayoutLMv3-B Table 0.979 #5 of 15 Archive leaderboard report
Document Layout Analysis PubLayNet val LayoutLMv3-B Text 0.945 #5 of 15 Archive leaderboard report
Document Layout Analysis PubLayNet val LayoutLMv3-B Title 0.906 #5 of 15 Archive leaderboard report
Entity Linking EC-FUNSD LayoutLMv3 (large) F1 78.14 #5 of 8 Archive leaderboard report
Entity Linking EC-FUNSD LayoutLMv3 (base) F1 67.47 #8 of 8 Archive leaderboard report
Key Information Extraction CORD LayoutLMv3 Large F1 97.46 #3 of 9 Archive leaderboard report
Key Information Extraction EPHOIE LayoutLMv3 Average F1 99.21 #1 of 1 Archive leaderboard report
Key-value Pair Extraction RFUND-EN LayoutLMv3 key-value pair F1 57.66 #7 of 13 Archive leaderboard report
Key-value Pair Extraction SIBR LayoutLMv3_base_chinese key-value pair F1 73.51 #4 of 7 Archive leaderboard report
Named Entity Recognition (NER) CORD-r LayoutLMv3 F1 82.72 #3 of 4 Archive leaderboard report
Named Entity Recognition (NER) FUNSD-r LayoutLMv3 F1 78.77 #2 of 4 Archive leaderboard report
Relation Extraction FUNSD LayoutLMv3 large F1 80.35 #4 of 9 Archive leaderboard report
Semantic entity labeling EC-FUNSD LayoutLMv3 (large) F1 83.88 #4 of 8 Archive leaderboard report
Semantic entity labeling EC-FUNSD LayoutLMv3 (base) F1 82.30 #8 of 8 Archive leaderboard report
Semantic entity labeling FUNSD LayoutLMv3 Large F1 92.08 #5 of 15 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections