Papers › CDistNet: Perceiving Multi-Domain Character Distance for Robust Text Recognition

CDistNet: Perceiving Multi-Domain Character Distance for Robust Text Recognition

22 Nov 2021arXiv:2111.11011archive 2025-07-28

Tianlun Zheng, Zhineng Chen, Shancheng Fang, Hongtao Xie, Yu-Gang Jiang

The Transformer-based encoder-decoder framework is becoming popular in scene text recognition, largely because it naturally integrates recognition clues from both visual and semantic domains. However, recent studies show that the two kinds of clues are not always well registered and therefore, feature and character might be misaligned in difficult text (e.g., with a rare shape). As a result, constraints such as character position are introduced to alleviate this problem. Despite certain success, visual and semantic are still separately modeled and they are merely loosely associated. In this paper, we propose a novel module called Multi-Domain Character Distance Perception (MDCDP) to establish a visually and semantically related position embedding. MDCDP uses the position embedding to query both visual and semantic features following the cross-attention mechanism. The two kinds of clues are fused into the position branch, generating a content-aware embedding that well perceives character spacing and orientation variants, character semantic affinities, and clues tying the two kinds of information. They are summarized as the multi-domain character distance. We develop CDistNet that stacks multiple MDCDPs to guide a gradually precise distance modeling. Thus, the feature-character alignment is well built even various recognition difficulties are presented. We verify CDistNet on ten challenging public datasets and two series of augmented datasets created by ourselves. The experiments demonstrate that CDistNet performs highly competitively. It not only ranks top-tier in standard benchmarks, but also outperforms recent popular methods by obvious margins on real and augmented datasets presenting severe text deformation, poor linguistic support, and rare character layouts. Code is available at https://github.com/simplify23/CDistNet.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

simplify23/CDistNet officialmentioned in papermentioned on GitHubpytorch report
chibohe/CdistNet-pytorch mentioned on GitHubpytorch report
topdu/openocr pytorchApache-2.0 report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Scene Text Recognition

1 archive task tag without a task page not shown.

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Scene Text Recognition CUTE80 CDistNet (Ours) Accuracy 89.58 #18 of 18 Archive leaderboard report
Scene Text Recognition ICDAR2013 CDistNet (Ours) Accuracy 97.67 #14 of 38 Archive leaderboard report
Scene Text Recognition ICDAR2015 CDistNet (Ours) Accuracy 86.25 #12 of 27 Archive leaderboard report
Scene Text Recognition IIIT5k CDistNet (Ours) Accuracy 96.57 #16 of 17 Archive leaderboard report
Scene Text Recognition SVT CDistNet (Ours) Accuracy 93.82 #19 of 37 Archive leaderboard report
Scene Text Recognition SVTP CDistNet (Ours) Accuracy 89.77 #15 of 17 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections