Datasets › gRefCOCO

gRefCOCO

1 Jan 2023 archive 2025-07-28

gRefCOCO is the first large-scale Generalized Referring Expression Segmentation dataset that contains multi-target, no-target, and single-target expressions.

Benchmarks archive 2025-07-28

All 2 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

Papers archive 2025-07-28

15 shown of 15 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 43. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
DeRIS: Decoupling Perception and Cognition for Enhanced Referring Image Segmentation through Loopback Synergy 1 1 2 Jul 2025 not harvested
SimVG: A Simple Framework for Visual Grounding with Decoupled Multi-modal Fusion 1 1 26 Sep 2024 ran 7 of 8 samples (1 unverified)
CoHD: A Counting-Aware Hierarchical Decoding Framework for Generalized Referring Expression Segmentation 2 1 24 May 2024 not harvested
Bring Adaptive Binding Prototypes to Generalized Referring Expression Segmentation 1 1 24 May 2024 not harvested
GROUNDHOG: Grounding Large Language Models to Holistic Segmentation 0 1 26 Feb 2024 not harvested
GSVA: Generalized Segmentation via Multimodal Large Language Models 1 3 15 Dec 2023 ran 2 of 10 samples (8 unverified)
GRES: Generalized Referring Expression Segmentation 2 1 1 Jun 2023 ran 2 of 5 samples (3 unverified)
Universal Instance Perception as Object Discovery and Retrieval 1 1 12 Mar 2023 ran 3 of 4 samples (1 unverified)
LAVT: Language-Aware Vision Transformer for Referring Image Segmentation 1 1 4 Dec 2021 not harvested
CRIS: CLIP-Driven Referring Image Segmentation 1 1 30 Nov 2021 ran 7 of 12 samples (5 unverified)
Vision-Language Transformer and Query Generation for Referring Segmentation 1 2 12 Aug 2021 not harvested
MDETR -- Modulated Detection for End-to-End Multi-Modal Understanding 5 1 26 Apr 2021 ran 6 of 11 samples (5 unverified)
Locate then Segment: A Strong Pipeline for Referring Image Segmentation 0 1 30 Mar 2021 not harvested
Multi-task Collaborative Network for Joint Referring Expression Comprehension and Segmentation 2 1 19 Mar 2020 ran 0 of 13 samples (13 unverified)
MAttNet: Modular Attention Network for Referring Expression Comprehension 1 1 24 Jan 2018 not harvested

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • gRefCOCO

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections