Papers › Multi-modal Text Recognition Networks: Interactive Enhancements between Visual and...

Multi-modal Text Recognition Networks: Interactive Enhancements between Visual and Semantic Features

30 Nov 2021arXiv:2111.15263archive 2025-07-28

Byeonghu Na, Yoonsik Kim, Sungrae Park

Linguistic knowledge has brought great benefits to scene text recognition by providing semantics to refine character sequences. However, since linguistic knowledge has been applied individually on the output sequence, previous methods have not fully utilized the semantics to understand visual clues for text recognition. This paper introduces a novel method, called Multi-modAl Text Recognition Network (MATRN), that enables interactions between visual and semantic features for better recognition performances. Specifically, MATRN identifies visual and semantic feature pairs and encodes spatial information into semantic features. Based on the spatial encoding, visual and semantic features are enhanced by referring to related features in the other modality. Furthermore, MATRN stimulates combining semantic features into visual features by hiding visual clues related to the character in the training phase. Our experiments demonstrate that MATRN achieves state-of-the-art performances on seven benchmarks with large margins, while naive combinations of two modalities show less-effective improvements. Further ablative studies prove the effectiveness of our proposed components. Our implementation is available at https://github.com/wp03052/MATRN.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

wp03052/MATRN officialmentioned in papermentioned on GitHubpytorchMIT report
byeonghu-na/matrn mentioned on GitHubpytorch report
topdu/openocr pytorchApache-2.0 report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Scene Text Recognition

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Scene Text Recognition CUTE80 MATRN Accuracy 93.5 #13 of 18 Archive leaderboard report
Scene Text Recognition ICDAR2013 MATRN Accuracy 97.9 #10 of 38 Archive leaderboard report
Scene Text Recognition ICDAR2015 MATRN Accuracy 86.6 #11 of 27 Archive leaderboard report
Scene Text Recognition IIIT5k MATRN Accuracy 96.6 #15 of 17 Archive leaderboard report
Scene Text Recognition SVT MATRN Accuracy 95 #15 of 37 Archive leaderboard report
Scene Text Recognition SVTP MATRN Accuracy 90.6 #13 of 17 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections