Papers › LayoutXLM: Multimodal Pre-training for Multilingual Visually-rich Document Understanding

LayoutXLM: Multimodal Pre-training for Multilingual Visually-rich Document Understanding

18 Apr 2021arXiv:2104.08836archive 2025-07-28

Yiheng Xu, Tengchao Lv, Lei Cui, Guoxin Wang, Yijuan Lu, Dinei Florencio, Cha Zhang, Furu Wei

Multimodal pre-training with text, layout, and image has achieved SOTA performance for visually-rich document understanding tasks recently, which demonstrates the great potential for joint learning across different modalities. In this paper, we present LayoutXLM, a multimodal pre-trained model for multilingual document understanding, which aims to bridge the language barriers for visually-rich document understanding. To accurately evaluate LayoutXLM, we also introduce a multilingual form understanding benchmark dataset named XFUND, which includes form understanding samples in 7 languages (Chinese, Japanese, Spanish, French, Italian, German, Portuguese), and key-value pairs are manually labeled for each language. Experiment results show that the LayoutXLM model has significantly outperformed the existing SOTA cross-lingual pre-trained models on the XFUND dataset. The pre-trained LayoutXLM model and the XFUND dataset are publicly available at https://aka.ms/layoutxlm.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

microsoft/unilm officialpytorch report
facebookresearch/data2vec_vision mentioned on GitHubpytorch report
huggingface/transformers mentioned on GitHubpytorch report
PaddlePaddle/PaddleOCR paddleApache-2.0 report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Document Image ClassificationFormKey-value Pair Extractiondocument understanding

Datasets

Introduced by this paper, per the archive.

XFUND

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Document Image Classification RVL-CDIP LayoutXLM Accuracy 95.21% #13 of 31 Archive leaderboard report
Key-value Pair Extraction RFUND-EN LayoutXLM_base key-value pair F1 53.98 #9 of 13 Archive leaderboard report
Key-value Pair Extraction SIBR LayoutXLM key-value pair F1 70.45 #6 of 7 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections