{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/layoutlm-pre-training-of-text-and-layout-for","title":"LayoutLM: Pre-training of Text and Layout for Document Image Understanding","arxiv_id":"1912.13318","date":"2019-12-31","proceeding":null,"authors":["Yiheng Xu","Minghao Li","Lei Cui","Shaohan Huang","Furu Wei","Ming Zhou"],"abstract":"Pre-training techniques have been verified successfully in a variety of NLP tasks in recent years. Despite the widespread use of pre-training models for NLP applications, they almost exclusively focus on text-level manipulation, while neglecting layout and style information that is vital for document image understanding. In this paper, we propose the \\textbf{LayoutLM} to jointly model interactions between text and layout information across scanned document images, which is beneficial for a great number of real-world document image understanding tasks such as information extraction from scanned documents. Furthermore, we also leverage image features to incorporate words' visual information into LayoutLM. To the best of our knowledge, this is the first time that text and layout are jointly learned in a single framework for document-level pre-training. It achieves new state-of-the-art results in several downstream tasks, including form understanding (from 70.72 to 79.27), receipt understanding (from 94.02 to 95.24) and document image classification (from 93.07 to 94.42). The code and pre-trained LayoutLM models are publicly available at \\url{https://aka.ms/layoutlm}.","url_abs":"https://arxiv.org/abs/1912.13318v5","url_pdf":"https://arxiv.org/pdf/1912.13318v5.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"layoutlm-pre-training-of-text-and-layout-for","repo_url":"https://github.com/microsoft/unilm","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"layoutlm-pre-training-of-text-and-layout-for","repo_url":"https://github.com/BordiaS/layoutlm","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"layoutlm-pre-training-of-text-and-layout-for","repo_url":"https://github.com/cydal/LayoutML_pytorch","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"layoutlm-pre-training-of-text-and-layout-for","repo_url":"https://github.com/doc-analysis/DocBank","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"Apache-2.0"}},{"paper_slug":"layoutlm-pre-training-of-text-and-layout-for","repo_url":"https://github.com/facebookresearch/data2vec_vision","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"layoutlm-pre-training-of-text-and-layout-for","repo_url":"https://github.com/huggingface/transformers","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"layoutlm-pre-training-of-text-and-layout-for","repo_url":"https://github.com/impira/docquery","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"layoutlm-pre-training-of-text-and-layout-for","repo_url":"https://github.com/kenAlparslan/Texttract","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}},{"paper_slug":"layoutlm-pre-training-of-text-and-layout-for","repo_url":"https://github.com/lulia0228/Document_IE","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null},{"paper_slug":"layoutlm-pre-training-of-text-and-layout-for","repo_url":"https://github.com/microsoft/unilm/tree/master/layoutlm","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"layoutlm-pre-training-of-text-and-layout-for","repo_url":"https://github.com/omarsou/layoutlm_CORD","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"layoutlm-pre-training-of-text-and-layout-for","repo_url":"https://github.com/thibaultdouzon/long-range-document-transformer","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"layoutlm-pre-training-of-text-and-layout-for","repo_url":"https://github.com/MS-P3/code3/tree/main/layoutlm","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null},{"paper_slug":"layoutlm-pre-training-of-text-and-layout-for","repo_url":"https://github.com/MindSpore-scientific-2/code-14/tree/main/layoutlm","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null},{"paper_slug":"layoutlm-pre-training-of-text-and-layout-for","repo_url":"https://github.com/PaddlePaddle/PaddleNLP/tree/develop/paddlenlp/transformers/layoutlm","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"paddle","reach":null},{"paper_slug":"layoutlm-pre-training-of-text-and-layout-for","repo_url":"https://github.com/PaddlePaddle/PaddleOCR","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"paddle","reach":{"status":"ok","spdx":"Apache-2.0"}},{"paper_slug":"layoutlm-pre-training-of-text-and-layout-for","repo_url":"https://github.com/nunenuh/layoutlm.pytorch","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":null},{"paper_slug":"layoutlm-pre-training-of-text-and-layout-for","repo_url":"https://github.com/pwc-1/Paper-9/tree/main/layoutlm","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null},{"paper_slug":"layoutlm-pre-training-of-text-and-layout-for","repo_url":"https://github.com/yangyucheng000/papercode-2/tree/main/layout-diffusion-mindspore","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null}],"tasks":[{"task_slug":"document-ai","task_name":"Document AI"},{"task_slug":"document-image-classification","task_name":"Document Image Classification"},{"task_slug":"document-layout-analysis","task_name":"Document Layout Analysis"},{"task_slug":"image-classification","task_name":"Image Classification"},{"task_slug":"key-information-extraction","task_name":"Key Information Extraction"},{"task_slug":"relation-extraction","task_name":"Relation Extraction"},{"task_slug":"document-image-classification","task_name":"document-image-classification"},{"task_slug":"image-classification","task_name":"image-classification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/document-image-classification-on-rvl-cdip","task":"Document Image Classification","dataset":"RVL-CDIP","model":"Pre-trained LayoutLM","rank_in_archive_order":15,"of":31,"metrics":{"Accuracy":"94.42%","Parameters":"160M"},"uses_additional_data":false},{"leaderboard":"/sota/relation-extraction-on-funsd","task":"Relation Extraction","dataset":"FUNSD","model":"LayoutLM","rank_in_archive_order":9,"of":9,"metrics":{"F1":"42.83"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/1912.13318","atlas_url":"https://app.syntology.ai/?focus=1912.13318","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1912.13318"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/huggingface/transformers","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/PaddlePaddle/PaddleOCR","reach":{"status":"ok","spdx":"Apache-2.0"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/thibaultdouzon/long-range-document-transformer","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/facebookresearch/data2vec_vision","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/cydal/LayoutML_pytorch","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/omarsou/layoutlm_CORD","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/lulia0228/Document_IE","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/yangyucheng000/papercode-2/tree/main/layout-diffusion-mindspore","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/impira/docquery","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/PaddlePaddle/PaddleNLP/tree/develop/paddlenlp/transformers/layoutlm","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/nunenuh/layoutlm.pytorch","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/kenAlparslan/Texttract","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/MS-P3/code3/tree/main/layoutlm","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/doc-analysis/DocBank","reach":{"status":"ok","spdx":"Apache-2.0"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/microsoft/unilm","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/MindSpore-scientific-2/code-14/tree/main/layoutlm","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/pwc-1/Paper-9/tree/main/layoutlm","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/microsoft/unilm/tree/master/layoutlm","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/BordiaS/layoutlm","reach":{"status":"ok"}}],"summary":{"ran_draft_wrong":2,"unverified":1},"by_repo_kind":{"listed":{"samples":2,"ran":2,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":1,"samples":[{"code_sha256_prefix":"f2858c5b1e4ffe3a","entry":"collate_fn","repo":"nunenuh/layoutlm.pytorch","repo_kind":"listed","path":"unilm/layoutlm/examples/seq_labeling/run_seq_labeling.py","file_url":"https://github.com/nunenuh/layoutlm.pytorch/blob/HEAD/unilm/layoutlm/examples/seq_labeling/run_seq_labeling.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"f2858c5b1e4ffe3a"}},{"code_sha256_prefix":"292d7b1eea96d516","entry":"get_labels","repo":"nunenuh/layoutlm.pytorch","repo_kind":"listed","path":"unilm/layoutlm/examples/seq_labeling/run_seq_labeling.py","file_url":"https://github.com/nunenuh/layoutlm.pytorch/blob/HEAD/unilm/layoutlm/examples/seq_labeling/run_seq_labeling.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"292d7b1eea96d516"}},{"code_sha256_prefix":"3c241ecfe3749a6d","entry":"simple_accuracy","repo":null,"repo_kind":null,"path":null,"file_url":null,"link_basis":"identical_code_first_harvested_elsewhere","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":null,"inline_ok":false,"mcp_get_code":{"code_sha256":"3c241ecfe3749a6d"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}