Datasets › RefCOCO

RefCOCO

Introduced by Sahar Kazemzadeh et al. in ReferItGame: Referring to Objects in Photographs of Natural Scenes1 Oct 2014 archive 2025-07-28

The RefCOCO dataset is a referring expression generation (REG) dataset used for tasks related to understanding natural language expressions that refer to specific objects in images. Here are the key details about RefCOCO:

  1. Collection Method: The dataset was collected using the ReferitGame, a two-player game. In this game, the first player views an image with a segmented target object and writes a natural language expression referring to that object. The second player sees only the image and the referring expression and must click on the corresponding object. If both players perform correctly, they earn points and switch roles; otherwise, they receive a new object and image for description.

  2. Dataset Variants: RefCOCO: Contains 142,209 refer expressions for 50,000 objects across 19,994 images. RefCOCO+: Includes 141,564 expressions for 49,856 objects in 19,992 images. RefCOCOg: This variant has 25,799 images, 95,010 referring expressions, and 49,822 object instances.

  3. Language and Restrictions: RefCOCO allows any type of language in the referring expressions. RefCOCO+ disallows location words in expressions to focus purely on appearance-based descriptions (e.g., "the man in the yellow polka-dotted shirt") rather than viewer-dependent descriptions (e.g., "the second man from the left").

These datasets serve as valuable resources for tasks like referring expression segmentation, comprehension, and visual grounding in computer vision research.

Benchmarks archive 2025-07-28

All 11 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Referring Expression Segmentation RefCoCo val DeRIS-L Overall IoU 85.41 DeRIS: Decoupling Perception and Cognition for Enhanced... Dmmm1997/DeRIS 37 Compare
Referring Expression Segmentation RefCOCO+ val MLCD-Seg-7B Overall IoU 79.4 Multi-label Cluster Discrimination for Visual... deepglint/unicom 33 Compare
Referring Expression Segmentation RefCOCO+ testA HyperSeg Overall IoU 83.5 HyperSeg: Towards Universal Visual Segmentation with... congvvc/HyperSeg 30 Compare
Referring Expression Segmentation RefCOCO+ test B MLCD-Seg-7B Overall IoU 75.6 Multi-label Cluster Discrimination for Visual... deepglint/unicom 30 Compare
Referring Expression Segmentation RefCOCO testA DeRIS-L Overall IoU 86.49 DeRIS: Decoupling Perception and Cognition for Enhanced... Dmmm1997/DeRIS 13 Compare
Referring Expression Segmentation RefCOCO testB HyperSeg Overall IoU 83.4 HyperSeg: Towards Universal Visual Segmentation with... congvvc/HyperSeg 13 Compare
Visual Grounding RefCOCO+ testA Florence-2-large-ft Accuracy (%) 95.3 Florence-2: Advancing a Unified Representation for a... retkowsky/florence-2 7 Compare
Visual Grounding RefCOCO+ test B Florence-2-large-ft Accuracy (%) 92.0 Florence-2: Advancing a Unified Representation for a... retkowsky/florence-2 6 Compare
Visual Grounding RefCOCO+ val Florence-2-large-ft Accuracy (%) 93.4 Florence-2: Advancing a Unified Representation for a... retkowsky/florence-2 6 Compare
Referring Expression Segmentation RefCOCO DETRIS IoU 81.0 Densely Connected Parameter-Efficient Tuning for... jiaqihuang01/detris 4 Compare
Visual Grounding RefCOCO testA HYDRA IoU 61.7 HYDRA: A Hyper Agent for Dynamic Compositional Visual Reasoning ControlNet/HYDRA 1 Compare

Papers archive 2025-07-28

30 shown of 43 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 439. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
DeRIS: Decoupling Perception and Cognition for Enhanced Referring Image Segmentation through Loopback Synergy 1 6 2 Jul 2025 not harvested
SegAgent: Exploring Pixel Understanding Capabilities in MLLMs by Imitating Human Annotator Trajectories 1 2 11 Mar 2025 not harvested
Densely Connected Parameter-Efficient Tuning for Referring Image Segmentation 1 7 15 Jan 2025 ran 7 of 17 samples (10 unverified)
Multi-task Visual Grounding with Coarse-to-Fine Consistency Constraints 1 6 12 Jan 2025 ran 5 of 13 samples (8 unverified)
MaskRIS: Semantic Distortion-aware Data Augmentation for Referring Image Segmentation 1 12 28 Nov 2024 not harvested
HyperSeg: Towards Universal Visual Segmentation with Large Language Model 1 6 26 Nov 2024 ran 7 of 17 samples (10 unverified)
Multi-label Cluster Discrimination for Visual Representation Learning 1 6 24 Jul 2024 ran 7 of 11 samples (4 unverified)
SafaRi:Adaptive Sequence Transformer for Weakly Supervised Referring Expression Segmentation 0 6 2 Jul 2024 not harvested
EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model 1 6 28 Jun 2024 ran 3 of 3 samples (0 unverified; 2 pointer-only for licence)
Vision-Aware Text Features in Referring Image Segmentation: From Object Understanding to Context Understanding 1 6 12 Apr 2024 not harvested
PSALM: Pixelwise SegmentAtion with Large Multi-Modal Model 1 1 21 Mar 2024 ran 2 of 7 samples (5 unverified)
HYDRA: A Hyper Agent for Dynamic Compositional Visual Reasoning 1 2 19 Mar 2024 ran 5 of 8 samples (3 unverified)
GROUNDHOG: Grounding Large Language Models to Holistic Segmentation 0 4 26 Feb 2024 not harvested
Mask Grounding for Referring Image Segmentation 1 6 19 Dec 2023 ran 10 of 13 samples (3 unverified; 13 pointer-only for licence)
General Object Foundation Model for Images and Videos at Scale 1 3 14 Dec 2023 ran 8 of 13 samples (5 unverified)
EVP: Enhanced Visual Perception using Inverse Multi-Attentive Feature Refinement and Regularized Image-Text Alignment 1 3 13 Dec 2023 not harvested
Universal Segmentation at Arbitrary Granularity with Language Instruction 2 7 4 Dec 2023 ran 13 of 16 samples (3 unverified)
Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks 1 3 10 Nov 2023 not harvested
Bridging Vision and Language Encoders: Parameter-Efficient Tuning for Referring Image Segmentation 1 2 21 Jul 2023 ran 8 of 16 samples (8 unverified)
Hierarchical Open-vocabulary Universal Image Segmentation 1 2 3 Jul 2023 ran 2 of 4 samples (2 unverified)
GRES: Generalized Referring Expression Segmentation 2 4 1 Jun 2023 ran 2 of 5 samples (3 unverified)
Universal Instance Perception as Object Discovery and Retrieval 1 4 12 Mar 2023 ran 3 of 4 samples (1 unverified)
Unleashing Text-to-Image Diffusion Models for Visual Perception 2 1 3 Mar 2023 ran 4 of 5 samples (1 unverified)
PolyFormer: Referring Image Segmentation as Sequential Polygon Generation 1 8 14 Feb 2023 ran 5 of 6 samples (1 unverified; 6 pointer-only for licence)
mPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and Video 4 3 1 Feb 2023 ran 9 of 19 samples (10 unverified)
Toward Building General Foundation Models for Language, Vision, and Vision-Language Understanding Tasks 1 3 12 Jan 2023 not harvested
X²-VLM: All-In-One Pre-trained Model For Vision-Language Tasks 2 6 22 Nov 2022 ran 2 of 6 samples (4 unverified; 6 pointer-only for licence)
VLT: Vision-Language Transformer and Query Generation for Referring Segmentation 1 4 28 Oct 2022 ran 0 of 6 samples (6 unverified)
SeqTR: A Simple yet Universal Network for Visual Grounding 3 2 30 Mar 2022 ran 1 of 4 samples (3 unverified)
LAVT: Language-Aware Vision Transformer for Referring Image Segmentation 1 3 4 Dec 2021 not harvested

The full list of 43 is in the JSON twin.

Dataset loaders archive 2025-07-28

2 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

Unknown

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • RefCOCOg-val
  • RefCOCOg-test
  • RefCoco+
  • RefCOCO+ test B
  • RefCOCO+ testA
  • RefCOCO+ val
  • RefCOCO testB
  • RefCOCO testA
  • RefCoCo val
  • RefCOCO

10 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections