Home › Datasets › task › Phrase Grounding

Phrase Grounding datasets

archive 2025-07-28

6 datasets carry the task tag "Phrase Grounding" (the task itself: Phrase Grounding), ordered by the archive's paper count. Page 1 of 1: 6 shown of 6. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Phrase Grounding datasets 1–6 of 6

Visual Genome contains Visual Question Answering data in a multi-choice setting.
1,256 papers · 15 benchmarks
The Flickr30k dataset contains 31,000 images collected from Flickr, together with 5 reference sentences provided by human annotators.
880 papers · 9 benchmarks
The Flickr30K Entities dataset is an extension to the Flickr30K dataset.
142 papers · 2 benchmarks
MS-CXR (Making the Most of Text Semantics to Improve Biomedical Vision-Language Processing)
The MS-CXR dataset provides 1162 image–sentence pairs of bounding boxes and corresponding phrases, collected across eight different cardiopulmonary radiological findings, with an approximately equal number of pairs for each finding.
32 papers · 0 benchmarks
G-VUE (General-purpose Visual Understanding Evaluation)
General-purpose Visual Understanding Evaluation (G-VUE) is a comprehensive benchmark covering the full spectrum of visual cognitive abilities with four functional domains -- Perceive, Ground, Reason, and Act.
2 papers · 0 benchmarks
VD-Ref is a dataset with ground-truth mappings from both noun phrases and pronouns to image regions.
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.