Papers › COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images

COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images

26 Jan 2016arXiv:1601.07140archive 2025-07-28

Andreas Veit, Tomas Matera, Lukas Neumann, Jiri Matas, Serge Belongie

This paper describes the COCO-Text dataset. In recent years large-scale datasets like SUN and Imagenet drove the advancement of scene understanding and object recognition. The goal of COCO-Text is to advance state-of-the-art in text detection and recognition in natural images. The dataset is based on the MS COCO dataset, which contains images of complex everyday scenes. The images were not collected with text in mind and thus contain a broad variety of text instances. To reflect the diversity of text in natural scenes, we annotate text with (a) location in terms of a bounding box, (b) fine-grained classification into machine printed text and handwritten text, (c) classification into legible and illegible text, (d) script of the text and (e) transcriptions of legible text. The dataset contains over 173k text annotations in over 63k images. We provide a statistical analysis of the accuracy of our annotations. In addition, we present an analysis of three leading state-of-the-art photo Optical Character Recognition (OCR) approaches on our dataset. While scene text detection and recognition enjoys strong advances in recent years, we identify significant shortcomings motivating future work.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

LinearPi/OCR_Chinese mentioned on GitHubtf report
OzHsu23/chineseocr mentioned on GitHubtf report
witcher425/CHINESEOCR mentioned on GitHubtf report
xiaofengShi/CHINESE-OCR mentioned on GitHubtf report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

DiversityGeneral ClassificationObject RecognitionOptical Character RecognitionOptical Character Recognition (OCR)Scene Text DetectionScene UnderstandingText Detection

Datasets

Introduced by this paper, per the archive.

COCO-Text

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections