Papers › VisualWordGrid: Information Extraction From Scanned Documents Using A Multimodal Approach

VisualWordGrid: Information Extraction From Scanned Documents Using A Multimodal Approach

5 Oct 2020arXiv:2010.02358archive 2025-07-28

Mohamed Kerroumi, Othmane Sayem, Aymen Shabou

We introduce a novel approach for scanned document representation to perform field extraction. It allows the simultaneous encoding of the textual, visual and layout information in a 3-axis tensor used as an input to a segmentation model. We improve the recent Chargrid and Wordgrid \cite{chargrid} models in several ways, first by taking into account the visual modality, then by boosting its robustness in regards to small datasets while keeping the inference time low. Our approach is tested on public and private document-image datasets, showing higher performances compared to the recent state-of-the-art methods.

PaperPDFConference PDF

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Document Layout Analysis

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Document Layout Analysis RVL-CDIP VisualWordGrid FAR 28.7 #1 of 1 Archive leaderboard report
Document Layout Analysis RVL-CDIP VisualWordGrid WAR 18.7 #1 of 1 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections