{"url":"/method/layoutlmv2","slug":"layoutlmv2","name":"LayoutLMv2","full_name":"LayoutLMv2","full_name_withheld":false,"description_markdown":"**LayoutLMv2** is an architecture and pre-training method for document understanding. The model is pre-trained with a great number of unlabeled scanned document images from the IIT-CDIP dataset, where some images in the text-image pairs are randomly replaced with another document image to make the model learn whether the image and OCR texts are correlated or not. Meanwhile, it also integrates a spatial-aware self-attention mechanism into the Transformer architecture, so that the model can fully understand the relative positional relationship among different text blocks.\r\n\r\nSpecifically, an enhanced Transformer architecture is used, i.e. a multi-modal Transformer asisthe backbone of LayoutLMv2. The multi-modal Transformer accepts inputs of three modalities: text, image, and layout. The input of each modality is converted to an embedding sequence and fused by the encoder. The model establishes deep interactions within and between modalities by leveraging the powerful Transformer layers.","description_state":"present","introduced_year":null,"introduced_by":{"title":null,"paper":null,"first_author":null,"n_authors":0,"url_abs":null,"archive_paper_url":null},"source":{"url":"https://arxiv.org/abs/2012.14740v4","title":"LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Natural Language Processing","area_id":"natural-language-processing","collection":"Document Understanding Models","url":"/methods/category/document-understanding-models","pwc_aliases":[]}],"n_papers_tagged":3,"archive_num_papers":null,"papers_newest_first":[{"paper":"/paper/peneo-unifying-line-extraction-line-grouping","title":"PEneo: Unifying Line Extraction, Line Grouping, and Entity Linking for End-to-end Document Pair Extraction","date":"2024-01-07","arxiv_id":"2401.03472","n_code_links":1,"syntology":null},{"paper":null,"title":"ProtoNER: Few shot Incremental Learning for Named Entity Recognition using Prototypical Networks","date":"2023-10-03","arxiv_id":"2310.02372","n_code_links":0,"syntology":null},{"paper":"/paper/layoutlmv2-multi-modal-pre-training-for","title":"LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding","date":"2020-12-29","arxiv_id":"2012.14740","n_code_links":9,"syntology":null}],"papers_shown":3,"tasks":[{"task":"/task/key-information-extraction","name":"Key Information Extraction","papers":2},{"task":"/task/key-value-pair-extraction","name":"Key-value Pair Extraction","papers":2},{"task":"/task/relation-extraction","name":"Relation Extraction","papers":2},{"task":"/task/semantic-entity-labeling","name":"Semantic entity labeling","papers":2},{"task":"/task/document-understanding","name":"document understanding","papers":2},{"task":"/task/document-image-classification","name":"Document Image Classification","papers":1},{"task":"/task/document-layout-analysis","name":"Document Layout Analysis","papers":1},{"task":"/task/incremental-learning","name":"Incremental Learning","papers":1},{"task":"/task/language-modeling","name":"Language Modeling","papers":1},{"task":"/task/language-modelling","name":"Language Modelling","papers":1},{"task":"/task/cg","name":"NER","papers":1},{"task":"/task/named-entity-recognition-1","name":"Named Entity Recognition","papers":1},{"task":"/task/named-entity-recognition-ner","name":"Named Entity Recognition (NER)","papers":1},{"task":"/task/synthetic-data-generation","name":"Synthetic Data Generation","papers":1},{"task":"/task/task-2","name":"Task 2","papers":1},{"task":"/task/visual-question-answering-1","name":"Visual Question Answering","papers":1},{"task":"/task/visual-question-answering","name":"Visual Question Answering (VQA)","papers":1},{"task":"/task/named-entity-recognition","name":"named-entity-recognition","papers":1}],"tasks_shown":18,"n_tasks":18,"usage_by_year":[{"year":"2020","papers":1},{"year":"2023","papers":1},{"year":"2024","papers":1}],"row_source":"embedded","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/layoutlmv2"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}