Papers › LiLT: A Simple yet Effective Language-Independent Layout Transformer for Structured...

LiLT: A Simple yet Effective Language-Independent Layout Transformer for Structured Document Understanding

28 Feb 2022ACL 2022 5arXiv:2202.13669archive 2025-07-28

Jiapeng Wang, Lianwen Jin, Kai Ding

Structured document understanding has attracted considerable attention and made significant progress recently, owing to its crucial role in intelligent document processing. However, most existing related models can only deal with the document data of specific language(s) (typically English) included in the pre-training collection, which is extremely limited. To address this issue, we propose a simple yet effective Language-independent Layout Transformer (LiLT) for structured document understanding. LiLT can be pre-trained on the structured documents of a single language and then directly fine-tuned on other languages with the corresponding off-the-shelf monolingual/multilingual pre-trained textual models. Experimental results on eight languages have shown that LiLT can achieve competitive or even superior performance on diverse widely-used downstream benchmarks, which enables language-independent benefit from the pre-training of document layout structure. Code and model are publicly available at https://github.com/jpWang/LiLT.

PaperPDFConference PDFCodeCode Syntology ran

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

For agents, Syntology's MCP tool lists every function and class Syntology harvested from this paper and whether it ran (how to connect): get_harvested_code_for_paper(arxiv_id="2202.13669")

Code

Syntology Ran 2 of 3 code samples harvested from 1 repository linked to this paper; 1 has no recorded run. Of those that ran: 1 ran · our draft was wrong; 1 ran with no contract checked.

By repository: official repository: 3 samples from 1 repository, 2 ran. The run record, sample by sample. “Ran” means executed on a synthesized input, not that the code is correct or reproduces the paper.

jpwang/lilt officialmentioned in papermentioned on GitHubpytorchMIT report
huggingface/transformers mentioned on GitHubpytorch report
MS-P3/code3 mindspore report
pwc-1/Paper-9 mindspore report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

3 samples harvested; 2 ran; 0 honoured the contract we drafted; 1 has no recorded run. Read from Syntology's graph 2026-09-24; that is when this build read the record, not when the samples ran.

1ran · our draft was wrong
1ran
1unverified

Licence: 0 of the 3 samples are pointer only, meaning Syntology does not serve that copy's text. This page shows no code text for any sample; each one links to its file in the repository.

Harvested from jpWang/LiLT. “Ran” means the sample executed on a synthesized input. It does not mean the output is correct, and nothing here reproduces the paper's results. “Honoured” and “violated” refer to a contract Syntology drafted from the code itself; “our draft was wrong” and “fixture could not drive it” are failures of Syntology's instrument, not of the code.

Each sample ends with its code_sha256, Syntology's identity for that exact code. An agent fetches the stored sample with Syntology's MCP tool get_code(code_sha256="…") (how to connect); click an identity to copy that call.

Repository labels, per sample. official repository: The archive marks this repository official for the paper. named in the paper: The archive records that the paper mentions this repository; it is not marked official. community (archive-listed): In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper. found in paper text by Syntology: Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted. community: Not in the archive's code links for this paper; a community repository Syntology harvested. Samples from a repository marked official are listed first. Licence labels name the repository's licence as recorded at harvest. “Pointer only” means Syntology does not serve that copy's text, for one of four reasons: no licence file was found; the licence was not identified; the licence is recorded as permissive but that copy's record is not marked cleared; or the licence is outside the permissive list Syntology serves text under (MIT, Apache-2.0, BSD and similar). Some licences outside that list permit redistribution, such as WTFPL, and GPL-3.0 under its conditions; they are simply not on the list. Hover a licence label for the reason. File links open the file on GitHub at the default branch, which may have changed since the harvest.

create_position_ids_from_input_ids jpWang/LiLT/LiLTfinetune/models/LiLTRobertaLike/modeling_LiLTRobertaLike.py official repository ran · our draft was wrong fingerprinted MIT (permissive) · da1c54dffed5d60c · report
get_last_checkpoint jpWang/LiLT/LiLTfinetune/evaluation.py official repository ran MIT (permissive) · 9f1130d32c6430f8 · report
re_score jpWang/LiLT/LiLTfinetune/evaluation.py official repository unverified MIT (permissive) · e12a52020d4a022c · report

Tasks

Document Image ClassificationKey Information ExtractionKey-value Pair ExtractionSemantic entity labelingdocument understanding

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Document Image Classification RVL-CDIP LiLT[EN-R]BASE Accuracy 95.68% #5 of 31 Archive leaderboard report
Key Information Extraction CORD LILT F1 96.07 #7 of 9 Archive leaderboard report
Key-value Pair Extraction RFUND-EN LiLT ([EN-R]_base) key-value pair F1 54.33 #8 of 13 Archive leaderboard report
Key-value Pair Extraction RFUND-EN LiLT ([InfoXLM]_base) key-value pair F1 52.18 #10 of 13 Archive leaderboard report
Key-value Pair Extraction SIBR LiLT ([InfoXLM]_base) key-value pair F1 72.76 #5 of 7 Archive leaderboard report
Semantic entity labeling FUNSD LILT F1 88.41 #10 of 15 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Absolute Position EncodingsAdamAttentionBPEDense ConnectionsDropoutLabel SmoothingLayer NormalizationLinear LayerMulti-Head AttentionPosition-Wise Feed-Forward LayerResidual ConnectionSoftmaxTransformer

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections