Papers › DTrOCR: Decoder-only Transformer for Optical Character Recognition

DTrOCR: Decoder-only Transformer for Optical Character Recognition

30 Aug 2023arXiv:2308.15996archive 2025-07-28

Masato Fujitake

Typical text recognition methods rely on an encoder-decoder structure, in which the encoder extracts features from an image, and the decoder produces recognized text from these features. In this study, we propose a simpler and more effective method for text recognition, known as the Decoder-only Transformer for Optical Character Recognition (DTrOCR). This method uses a decoder-only Transformer to take advantage of a generative language model that is pre-trained on a large corpus. We examined whether a generative language model that has been successful in natural language processing can also be effective for text recognition in computer vision. Our experiments demonstrated that DTrOCR outperforms current state-of-the-art methods by a large margin in the recognition of printed, handwritten, and scene text in both English and Chinese.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

DecoderHandwritten Text RecognitionLanguage ModelingLanguage ModellingOptical Character RecognitionOptical Character Recognition (OCR)Scene Text RecognitionTask 2

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Handwritten Text Recognition IAM DTrOCR 105M CER 2.38 #1 of 17 Archive leaderboard report
Optical Character Recognition (OCR) Benchmarking Chinese Text Recognition: Datasets, Baselines, and an Empirical Study DTrOCR Accuracy (%) 89.6 #1 of 7 Archive leaderboard report
Optical Character Recognition (OCR) Benchmarking Chinese Text Recognition: Datasets, Baselines, and an Empirical Study DTrOCR 105M Accuracy (%) 89.6 #2 of 7 Archive leaderboard report
Scene Text Recognition CUTE80 DTrOCR 105M Accuracy 99.1 #6 of 18 Archive leaderboard report
Scene Text Recognition ICDAR2013 DTrOCR 105M Accuracy 99.4 #2 of 38 Archive leaderboard report
Scene Text Recognition ICDAR2015 DTrOCR 105M Accuracy 93.5 #1 of 27 Archive leaderboard report
Scene Text Recognition IIIT5k DTrOCR 105M Accuracy 99.6 #2 of 17 Archive leaderboard report
Scene Text Recognition SVT DTrOCR 105M Accuracy 98.9 #2 of 37 Archive leaderboard report
Scene Text Recognition SVTP DTrOCR 105M Accuracy 98.6 #1 of 17 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Absolute Position EncodingsAdamAttentionBPEDense ConnectionsDropoutLabel SmoothingLayer NormalizationLinear LayerMulti-Head AttentionPosition-Wise Feed-Forward LayerResidual ConnectionSoftmaxTransformer

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections