Methods › Natural Language Processing › Document Understanding Models › LayoutLMv2

LayoutLMv2

3 papers tagged archive 2025-07-28

archive 2025-07-28 Description, source and code snippet are the archive's method entry.

LayoutLMv2 is an architecture and pre-training method for document understanding. The model is pre-trained with a great number of unlabeled scanned document images from the IIT-CDIP dataset, where some images in the text-image pairs are randomly replaced with another document image to make the model learn whether the image and OCR texts are correlated or not. Meanwhile, it also integrates a spatial-aware self-attention mechanism into the Transformer architecture, so that the model can fully understand the relative positional relationship among different text blocks.

Specifically, an enhanced Transformer architecture is used, i.e. a multi-modal Transformer asisthe backbone of LayoutLMv2. The multi-modal Transformer accepts inputs of three modalities: text, image, and layout. The input of each modality is converted to an embedding sequence and fused by the encoder. The model establishes deep interactions within and between modalities by leveraging the powerful Transformer layers.

Source: LayoutLMv2: Multi-modal Pre-training for Visually-Rich...

Papers archive 2025-07-28

3 shown of 3, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.

Tasks archive 2025-07-28

18 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.

TaskPapers
Key Information Extraction2
Key-value Pair Extraction2
Relation Extraction2
Semantic entity labeling2
document understanding2
Document Image Classification1
Document Layout Analysis1
Incremental Learning1
Language Modeling1
Language Modelling1
NER1
Named Entity Recognition1
Named Entity Recognition (NER)1
Synthetic Data Generation1
Task 21
Visual Question Answering1
Visual Question Answering (VQA)1
named-entity-recognition1

Usage over time archive 2025-07-28

Papers per year tagged with LayoutLMv2: 2020 to 2024, peak 1 1 0 2020: 1 paper 2020 2021: 0 papers 2021 2022: 0 papers 2022 2023: 1 paper 2023 2024: 1 paper 2024
Papers per year the archive tags with this method, by the paper's archive date (3 dated). Bars are counts, not a trend claim.

Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).

Categories archive 2025-07-28

Document Understanding Models

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections