Datasets › NoCaps

NoCaps

Introduced by Harsh Agrawal et al. in nocaps: novel object captioning at scale archive 2025-07-28

The nocaps benchmark consists of 166,100 human-generated captions describing 15,100 images from the OpenImages validation and test sets.

Source: nocaps: novel object captioning at scale Image Source: https://nocaps.org/

Benchmarks archive 2025-07-28

All 13 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Image Captioning nocaps in-domain PaLI CIDEr 149.1 PaLI: A Jointly-Scaled Multilingual Language-Image Model google-research/big_vision 41 Compare
Image Captioning nocaps near-domain GIT2, Single Model CIDEr 125.51 GIT: A Generative Image-to-text Transformer for Vision... microsoft/GenerativeImage2Text 40 Compare
Image Captioning nocaps out-of-domain PaLI CIDEr 126.67 PaLI: A Jointly-Scaled Multilingual Language-Image Model google-research/big_vision 40 Compare
Image Captioning nocaps entire Lyrics CIDEr 126.8 Lyrics: Boosting Fine-grained Language-Vision Alignment... — 39 Compare
Image Captioning nocaps-XD entire GIT2 CIDEr 124.77 GIT: A Generative Image-to-text Transformer for Vision... microsoft/GenerativeImage2Text 12 Compare
Image Captioning nocaps-val-in-domain BLIP-2 ViT-G FlanT5 XL (zero-shot) CIDEr 123.7 BLIP-2: Bootstrapping Language-Image Pre-training with... huggingface/transformers +16 11 Compare
Image Captioning nocaps-val-overall BLIP-2 ViT-G FlanT5 XL (zero-shot) CIDEr 121.6 BLIP-2: Bootstrapping Language-Image Pre-training with... huggingface/transformers +16 11 Compare
Image Captioning nocaps-XD in-domain GIT2 CIDEr 124.18 GIT: A Generative Image-to-text Transformer for Vision... microsoft/GenerativeImage2Text 11 Compare
Image Captioning nocaps-XD near-domain GIT2 CIDEr 125.51 GIT: A Generative Image-to-text Transformer for Vision... microsoft/GenerativeImage2Text 11 Compare
Image Captioning nocaps-XD out-of-domain GIT2 CIDEr 122.27 GIT: A Generative Image-to-text Transformer for Vision... microsoft/GenerativeImage2Text 11 Compare
Image Captioning nocaps-val-near-domain BLIP-2 ViT-G FlanT5 XL (zero-shot) CIDEr 120.2 BLIP-2: Bootstrapping Language-Image Pre-training with... huggingface/transformers +16 10 Compare
Image Captioning nocaps-val-out-domain BLIP-2 ViT-G FlanT5 XL (zero-shot) CIDEr 124.8 BLIP-2: Bootstrapping Language-Image Pre-training with... huggingface/transformers +16 10 Compare
Image Captioning nocaps val Prismer CIDEr 107.9 Prismer: A Vision-Language Model with Multi-Task Experts nvlabs/prismer +1 3 Compare

Papers archive 2025-07-28

17 shown of 17 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 175. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Lyrics: Boosting Fine-grained Language-Vision Alignment and Comprehension via Semantic-aware Visual Objects 0 1 8 Dec 2023 not harvested
Prismer: A Vision-Language Model with Multi-Task Experts 2 2 4 Mar 2023 not harvested
BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models 17 12 30 Jan 2023 ran 4 of 8 samples (4 unverified; 1 pointer-only for licence)
OmniVL:One Foundation Model for Image-Language and Video-Language Tasks 0 4 15 Sep 2022 not harvested
PaLI: A Jointly-Scaled Multilingual Language-Image Model 1 5 14 Sep 2022 ran 2 of 4 samples (2 unverified)
GRIT: Faster and Better Image captioning Transformer Using Dual Visual Features 2 2 20 Jul 2022 not harvested
Language Models are General-Purpose Interfaces 1 1 13 Jun 2022 not harvested
GIT: A Generative Image-to-text Transformer for Vision and Language 1 15 27 May 2022 ran 0 of 21 samples (21 unverified)
BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation 9 8 28 Jan 2022 not harvested
Scaling Up Vision-Language Pre-training for Image Captioning 0 6 24 Nov 2021 not harvested
ClipCap: CLIP Prefix for Image Captioning 4 8 18 Nov 2021 ran 3 of 6 samples (3 unverified)
SimVLM: Simple Visual Language Model Pretraining with Weak Supervision 2 8 24 Aug 2021 ran 18 of 37 samples (19 unverified; 28 pointer-only for licence)
Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts 3 4 17 Feb 2021 not harvested
Unifying Vision-and-Language Tasks via Text Generation 2 1 4 Feb 2021 ran 2 of 12 samples (10 unverified)
VinVL: Revisiting Visual Representations in Vision-Language Models 7 8 2 Jan 2021 ran 2 of 2 samples (0 unverified; 1 pointer-only for licence)
VIVO: Visual Vocabulary Pre-Training for Novel Object Captioning 0 8 28 Sep 2020 not harvested
Oscar: Object-Semantics Aligned Pre-training for Vision-Language Tasks 4 1 13 Apr 2020 ran 11 of 23 samples (12 unverified)

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

Unknown

Modalities archive 2025-07-28

No modality tagged.

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • nocaps val
  • nocaps-val-overall
  • nocaps-val-out-domain
  • nocaps-val-near-domain
  • nocaps-val-in-domain
  • nocaps-XD out-of-domain
  • nocaps-XD near-domain
  • nocaps-XD in-domain
  • nocaps-XD entire
  • nocaps out-of-domain
  • nocaps near-domain
  • nocaps in-domain
  • nocaps entire
  • NoCaps

14 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections