Papers › DUBLIN -- Document Understanding By Language-Image Network

DUBLIN -- Document Understanding By Language-Image Network

23 May 2023arXiv:2305.14218archive 2025-07-28

Kriti Aggarwal, Aditi Khandelwal, Kumar Tanmay, Owais Mohammed Khan, Qiang Liu, Monojit Choudhury, Hardik Hansrajbhai Chauhan, Subhojit Som, Vishrav Chaudhary, Saurabh Tiwary

Visual document understanding is a complex task that involves analyzing both the text and the visual elements in document images. Existing models often rely on manual feature engineering or domain-specific pipelines, which limit their generalization ability across different document types and languages. In this paper, we propose DUBLIN, which is pretrained on web pages using three novel objectives: Masked Document Text Generation Task, Bounding Box Task, and Rendered Question Answering Task, that leverage both the spatial and semantic information in the document images. Our model achieves competitive or state-of-the-art results on several benchmarks, such as Web-Based Structural Reading Comprehension, Document Visual Question Answering, Key Information Extraction, Diagram Understanding, and Table Question Answering. In particular, we show that DUBLIN is the first pixel-based model to achieve an EM of 77.75 and F1 of 84.25 on the WebSRC dataset. We also show that our model outperforms the current pixel-based SOTA models on DocVQA, InfographicsVQA, OCR-VQA and AI2D datasets by 4.6%, 6.5%, 2.6% and 21%, respectively. We also achieve competitive performance on RVL-CDIP document classification. Moreover, we create new baselines for text-based datasets by rendering them as document images to promote research in this direction.

PaperPDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Document ClassificationFeature EngineeringKey Information ExtractionOptical Character Recognition (OCR)Question AnsweringReading ComprehensionText GenerationVisual Question AnsweringVisual Question Answering (VQA)document understanding

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Visual Question Answering (VQA) AI2D DUBLIN EM 51.11 #4 of 4 Archive leaderboard report
Visual Question Answering (VQA) DeepForm DUBLIN F1 62.23 #1 of 1 Archive leaderboard report
Visual Question Answering (VQA) DocVQA test DUBLIN (variable resolution) ANLS 0.803 #22 of 33 Archive leaderboard report
Visual Question Answering (VQA) DocVQA test DUBLIN ANLS 0.782 #24 of 33 Archive leaderboard report
Visual Question Answering (VQA) InfographicVQA DUBLIN (variable resolution) ANLS 42.6 #17 of 21 Archive leaderboard report
Visual Question Answering (VQA) InfographicVQA DUBLIN ANLS 36.82 #21 of 21 Archive leaderboard report
Visual Question Answering (VQA) WebSRC DUBLIN EM 77.75 #1 of 1 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections