Datasets › TextCaps

TextCaps

Introduced by Oleksii Sidorov et al. in TextCaps: a Dataset for Image Captioning with Reading Comprehension archive 2025-07-28

Contains 145k captions for 28k images. The dataset challenges a model to recognize text, relate it to its visual context, and decide what part of the text to copy or paraphrase, requiring spatial, semantic, and visual reasoning between multiple text tokens and visual entities, such as objects.

Source: TextCaps: a Dataset for Image Captioning with Reading Comprehension

Benchmarks archive 2025-07-28

All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Image Captioning TextCaps 2020 TAP CIDEr 103.22 — — 11 Compare

Papers archive 2025-07-28

No paper in the archive has a leaderboard row on this dataset; the archive counts 98 papers for it but never published that list.

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • TextCaps 2020
  • TextCaps

2 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections