Papers › DoPTA: Improving Document Layout Analysis using Patch-Text Alignment

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment

17 Dec 2024arXiv:2412.12902archive 2025-07-28

Nikitha SR, Tarun Ram Menta, Mausoom Sarkar

The advent of multimodal learning has brought a significant improvement in document AI. Documents are now treated as multimodal entities, incorporating both textual and visual information for downstream analysis. However, works in this space are often focused on the textual aspect, using the visual space as auxiliary information. While some works have explored pure vision based techniques for document image understanding, they require OCR identified text as input during inference, or do not align with text in their learning procedure. Therefore, we present a novel image-text alignment technique specially designed for leveraging the textual information in document images to improve performance on visual tasks. Our document encoder model DoPTA - trained with this technique demonstrates strong performance on a wide range of document image understanding tasks, without requiring OCR during inference. Combined with an auxiliary reconstruction objective, DoPTA consistently outperforms larger models, while using significantly lesser pre-training compute. DoPTA also sets new state-of-the art results on D4LA, and FUNSD, two challenging document visual analysis benchmarks.

PaperPDF

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Document AIDocument Image ClassificationDocument Layout AnalysisOptical Character Recognition (OCR)

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Document Image Classification RVL-CDIP DoPTA Accuracy 94.12% #16 of 31 Archive leaderboard report
Document Image Classification RVL-CDIP DoPTA Parameters 85M #16 of 31 Archive leaderboard report
Document Layout Analysis D4LA DoPTA mAP 70.72 #1 of 3 Archive leaderboard report
Document Layout Analysis D4LA DoPTA Model Parameters 85M #1 of 3 Archive leaderboard report
Document Layout Analysis PubLayNet val DoPTA Figure 0.970 #6 of 15 Archive leaderboard report
Document Layout Analysis PubLayNet val DoPTA List 0.957 #6 of 15 Archive leaderboard report
Document Layout Analysis PubLayNet val DoPTA Overall 0.949 #6 of 15 Archive leaderboard report
Document Layout Analysis PubLayNet val DoPTA Table 0.977 #6 of 15 Archive leaderboard report
Document Layout Analysis PubLayNet val DoPTA Text 0.944 #6 of 15 Archive leaderboard report
Document Layout Analysis PubLayNet val DoPTA Title 0.895 #6 of 15 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

ALIGN

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections