Datasets › Flickr30K Entities

Flickr30K Entities

Introduced by Bryan A. Plummer et al. in Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models1 Jan 2015 archive 2025-07-28

The Flickr30K Entities dataset is an extension to the Flickr30K dataset. It augments the original 158k captions with 244k coreference chains, linking mentions of the same entities across different captions for the same image, and associating them with 276k manually annotated bounding boxes. This is used to define a new benchmark for localization of textual entity mentions in an image.

Source: http://bryanplummer.com/Flickr30kEntities/ Image Source: http://bryanplummer.com/Flickr30kEntities/

Benchmarks archive 2025-07-28

All 2 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Phrase Grounding Flickr30k Entities Test GLIPv2 R@1 87.7 GLIPv2: Unifying Localization and Vision-Language Understanding microsoft/GLIP 18 Compare
Phrase Grounding Flickr30k Entities Dev Fiber-B R@1 87.1 Coarse-to-Fine Vision-Language Pre-training with Fusion... microsoft/fiber 3 Compare

Papers archive 2025-07-28

16 shown of 16 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 142. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Coarse-to-Fine Vision-Language Pre-training with Fusion in the Backbone 1 2 15 Jun 2022 ran 2 of 3 samples (1 unverified; 2 pointer-only for licence)
GLIPv2: Unifying Localization and Vision-Language Understanding 1 1 12 Jun 2022 ran 2 of 2 samples (0 unverified)
PEVL: Position-enhanced Pre-training and Prompt Tuning for Vision-language Models 1 2 23 May 2022 ran 3 of 7 samples (4 unverified)
Grounded Language-Image Pre-training 3 1 7 Dec 2021 ran 2 of 2 samples (0 unverified; 2 pointer-only for licence)
MDETR -- Modulated Detection for End-to-End Multi-Modal Understanding 5 1 26 Apr 2021 ran 6 of 11 samples (5 unverified)
Disentangled Motif-aware Graph Learning for Phrase Grounding 0 1 13 Apr 2021 not harvested
Learning Cross-modal Context Graph for Visual Grounding 1 1 13 Feb 2020 not harvested
Phrase Grounding by Soft-Label Chain Conditional Random Field 1 1 1 Sep 2019 not harvested
VisualBERT: A Simple and Performant Baseline for Vision and Language 10 2 9 Aug 2019 ran 4 of 9 samples (5 unverified; 6 pointer-only for licence)
Bilinear Attention Networks 8 1 21 May 2018 ran 4 of 13 samples (9 unverified)
Rethinking Diversified and Discriminative Proposal Generation for Visual Grounding 1 1 9 May 2018 not harvested
Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding 10 1 6 Jun 2016 not harvested
Learning Deep Structure-Preserving Image-Text Embeddings 0 1 19 Nov 2015 not harvested
Natural Language Object Retrieval 1 1 13 Nov 2015 ran 2 of 3 samples (1 unverified; 3 pointer-only for licence)
Grounding of Textual Phrases in Images by Reconstruction 3 1 12 Nov 2015 not harvested
Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models 2 3 19 May 2015 ran 2 of 2 samples (0 unverified; 2 pointer-only for licence)

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

Custom (research-only, non-commercial)

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • Flickr30k Entities Dev
  • Flickr30k Entities Test
  • Flickr30K Entities

3 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections