Datasets › Flickr30k

Flickr30k

Introduced by Peter Young et al. in From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions1 Jan 2014 archive 2025-07-28

The Flickr30k dataset contains 31,000 images collected from Flickr, together with 5 reference sentences provided by human annotators.

Source: Guiding Long-Short Term Memory for Image Caption Generation

Image Source: Dual-Path Convolutional Image-Text Embedding with Instance Loss

Benchmarks archive 2025-07-28

All 9 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Cross-Modal Retrieval Flickr30k X2-VLM (large) Image-to-text R@1 98.8 X²-VLM: All-In-One Pre-trained Model For Vision-Language Tasks zengyan-97/x-vlm +1 27 Compare
Zero-Shot Cross-Modal Retrieval Flickr30k InternVL-G Image-to-text R@1 95.7 InternVL: Scaling up Vision Foundation Models and... opengvlab/internvl +1 22 Compare
Image Retrieval Flickr30K 1K test X-VLM (base) R@1 86.9 Multi-Grained Vision Language Pre-Training: Aligning... zengyan-97/x-vlm 18 Compare
Image-to-Text Retrieval Flickr30k InternVL-G-FT (finetuned, w/o ranking) Recall@1 97.9 InternVL: Scaling up Vision Foundation Models and... opengvlab/internvl +1 11 Compare
Image Retrieval Flickr30k BLIP-2 ViT-G (zero-shot, 1K test set) Recall@10 98.9 BLIP-2: Bootstrapping Language-Image Pre-training with... huggingface/transformers +16 9 Compare
Node Classification Flickr GCN+GAugM (Zhao et al., 2021) Accuracy 0.682 Data Augmentation for Graph Neural Networks zhao-tong/GAug +1 8 Compare
Image Captioning Flickr30k Captions test Unified VLP BLEU-4 30.1 Unified Vision-Language Pre-Training for Image Captioning and VQA rmokady/clip_prefix_caption +2 7 Compare
Phrase Grounding Flickr30k GBS Ensemble + 12-in-1 Pointing Game Accuracy 85.9 Detector-Free Weakly Supervised Grounding by Separation aarbelle/GroundingBySeparation 3 Compare
Semi Supervised Learning for Image Captioning Flickr30k CapDec CIDEr 39.1 Text-Only Training for Image Captioning using Noise-Injected CLIP davidhuji/capdec +3 1 Compare

Papers archive 2025-07-28

30 shown of 69 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 880. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training 1 2 2 Dec 2024 not harvested
3SHNet: Boosting Image-Sentence Retrieval via Visual Semantic-Spatial Self-Highlighting 1 1 26 Apr 2024 ran 10 of 10 samples (0 unverified)
Dynamic Self-adaptive Multiscale Distillation from Pre-trained Multimodal Large Model for Efficient Cross-modal Representation Learning 1 1 16 Apr 2024 not harvested
M2-Encoder: Advancing Bilingual Image-Text Understanding by Large-scale Efficient Pretraining 1 1 29 Jan 2024 not harvested
InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks 2 4 21 Dec 2023 ran 2 of 2 samples (0 unverified; 2 pointer-only for licence)
Implicit Differentiable Outlier Detection Enable Robust Deep Multimodal Analysis 1 1 21 Sep 2023 not harvested
VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset 2 2 29 May 2023 ran 15 of 42 samples (27 unverified)
ONE-PEACE: Exploring One General Representation Model Toward Unlimited Modalities 2 1 18 May 2023 ran 2 of 7 samples (5 unverified)
Region-Aware Pretraining for Open-Vocabulary Object Detection with Vision Transformers 2 1 11 May 2023 ran 5 of 8 samples (3 unverified)
MaMMUT: A Simple Architecture for Joint Learning for MultiModal Tasks 1 1 29 Mar 2023 ran 3 of 3 samples (0 unverified)
Plug-and-Play Regulators for Image-Text Matching 1 2 23 Mar 2023 not harvested
BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models 17 4 30 Jan 2023 ran 4 of 8 samples (4 unverified; 1 pointer-only for licence)
HADA: A Graph-based Amalgamation Framework in Image-text Retrieval 2 5 11 Jan 2023 not harvested
NAPReg: Nouns As Proxies Regularization for Semantically Aware Cross-Modal Embeddings 1 1 7 Jan 2023 not harvested
Position-guided Text Prompt for Vision-Language Pre-training 1 1 19 Dec 2022 ran 3 of 5 samples (2 unverified)
Reproducible scaling laws for contrastive language-image learning 5 1 14 Dec 2022 ran 3 of 3 samples (0 unverified; 3 pointer-only for licence)
X²-VLM: All-In-One Pre-trained Model For Vision-Language Tasks 2 2 22 Nov 2022 ran 2 of 6 samples (4 unverified; 6 pointer-only for licence)
AltCLIP: Altering the Language Encoder in CLIP for Extended Language Capabilities 2 1 12 Nov 2022 ran 2 of 11 samples (9 unverified)
Text-Only Training for Image Captioning using Noise-Injected CLIP 4 1 1 Nov 2022 ran 4 of 5 samples (1 unverified; 1 pointer-only for licence)
Dissecting Deep Metric Learning Losses for Image-Text Retrieval 2 1 21 Oct 2022 not harvested
A Comprehensive Study on Large-Scale Graph Training: Benchmarking and Rethinking 2 1 14 Oct 2022 ran 4 of 6 samples (2 unverified)
ERNIE-ViL 2.0: Multi-view Contrastive Learning for Image-Text Pre-training 1 3 30 Sep 2022 not harvested
OmniVL:One Foundation Model for Image-Language and Video-Language Tasks 0 1 15 Sep 2022 not harvested
Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks 2 2 22 Aug 2022 not harvested
Language Models are General-Purpose Interfaces 1 1 13 Jun 2022 not harvested
CoCa: Contrastive Captioners are Image-Text Foundation Models 6 1 4 May 2022 ran 9 of 17 samples (8 unverified)
Flamingo: a Visual Language Model for Few-Shot Learning 5 1 29 Apr 2022 ran 18 of 24 samples (6 unverified; 7 pointer-only for licence)
ViSTA: Vision and Scene Text Aggregation for Cross-Modal Retrieval 0 1 31 Mar 2022 not harvested
Florence: A New Foundation Model for Computer Vision 2 1 22 Nov 2021 not harvested
Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts 1 2 16 Nov 2021 ran 1 of 1 samples (0 unverified)

The full list of 69 is in the JSON twin.

Dataset loaders archive 2025-07-28

3 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

Custom (research-only, non-commercial)

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • Flickr30k
  • Flickr
  • Flickr30k Captions test
  • Flickr30K 1K test

4 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections