Papers › GeoLayoutLM: Geometric Pre-training for Visual Information Extraction

GeoLayoutLM: Geometric Pre-training for Visual Information Extraction

21 Apr 2023CVPR 2023 1arXiv:2304.10759archive 2025-07-28

Chuwei Luo, Changxu Cheng, Qi Zheng, Cong Yao

Visual information extraction (VIE) plays an important role in Document Intelligence. Generally, it is divided into two tasks: semantic entity recognition (SER) and relation extraction (RE). Recently, pre-trained models for documents have achieved substantial progress in VIE, particularly in SER. However, most of the existing models learn the geometric representation in an implicit way, which has been found insufficient for the RE task since geometric information is especially crucial for RE. Moreover, we reveal another factor that limits the performance of RE lies in the objective gap between the pre-training phase and the fine-tuning phase for RE. To tackle these issues, we propose in this paper a multi-modal framework, named GeoLayoutLM, for VIE. GeoLayoutLM explicitly models the geometric relations in pre-training, which we call geometric pre-training. Geometric pre-training is achieved by three specially designed geometry-related pre-training tasks. Additionally, novel relation heads, which are pre-trained by the geometric pre-training tasks and fine-tuned for RE, are elaborately designed to enrich and enhance the feature representation. According to extensive experiments on standard VIE benchmarks, GeoLayoutLM achieves highly competitive scores in the SER task and significantly outperforms the previous state-of-the-arts for RE (\eg, the F1 score of RE on FUNSD is boosted from 80.35\% to 89.45\%). The code and models are publicly available at https://github.com/AlibabaResearch/AdvancedLiterateMachinery/tree/main/DocumentUnderstanding/GeoLayoutLM

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

alibabaresearch/advancedliteratemachinery officialmentioned in papermentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Document AIEntity LinkingKey Information ExtractionKey-value Pair ExtractionRelation ExtractionSemantic entity labelingentity_extraction

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Entity Linking FUNSD GeoLayoutLM F1 89.45 #1 of 7 Archive leaderboard report
Key Information Extraction CORD GeoLayoutLM F1 97.97 #2 of 9 Archive leaderboard report
Key-value Pair Extraction RFUND-EN GeoLayoutLM key-value pair F1 69.03 #6 of 13 Archive leaderboard report
Relation Extraction FUNSD GeoLayoutLM F1 89.45 #2 of 9 Archive leaderboard report
Relation Extraction FUNSD LayoutLMv3 large F1 80.35 #5 of 9 Archive leaderboard report
Semantic entity labeling FUNSD GeoLayoutLM F1 92.86 #4 of 15 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections