Datasets › Visual Genome

Visual Genome

Introduced by Ranjay Krishna et al. in Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations1 Jan 2017 archive 2025-07-28

Visual Genome contains Visual Question Answering data in a multi-choice setting. It consists of 101,174 images from MSCOCO with 1.7 million QA pairs, 17 questions per image on average. Compared to the Visual Question Answering dataset, Visual Genome represents a more balanced distribution over 6 question types: What, Where, When, Who, Why and How. The Visual Genome dataset also presents 108K images with densely annotated objects, attributes and relationships.

Source: RaAM: A Relation-aware Attention Model for Visual Question Answering Image Source: Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations

Benchmarks archive 2025-07-28

All 15 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Unbiased Scene Graph Generation Visual Genome IETrans (MOTIFS-ResNeXt-101-FPN backbone; PredCls mode) ng-mR@20 36.0 Fine-Grained Scene Graph Generation with Data Transfer waxnkw/ietrans-sgg.pytorch +1 31 Compare
Scene Graph Generation Visual Genome SpeaQ (without reweighting) Recall@50 32.9 Groupwise Query Specialization and Quality-Aware... mlvlab/speaq 19 Compare
Layout-to-Image Generation Visual Genome 128x128 LayoutDiffusion FID 16.35 LayoutDiffusion: Controllable Diffusion Model for... zgctroy/layoutdiffusion +1 5 Compare
Dense Captioning Visual Genome ControlCap mAP 18.2 ControlCap: Controllable Region-level Captioning callsys/controlcap 4 Compare
Layout-to-Image Generation Visual Genome 64x64 OC-GAN FID 20.27 Object-Centric Image Generation from Layouts — 4 Compare
Layout-to-Image Generation Visual Genome 256x256 LayoutDiffusion FID 15.63 LayoutDiffusion: Controllable Diffusion Model for... zgctroy/layoutdiffusion +1 4 Compare
Multi-label Image Recognition with Partial Labels Visual Genome DSRB Average mAP 46 Dual-Perspective Semantic-Aware Representation Blending... hcplab-sysu/hcp-mlr-pl 4 Compare
Object Detection Visual Genome KnowZRel MAP 44 KnowZRel: Common Sense Knowledge-based Zero-Shot... jaleedkhan/zsrr-sgg 4 Compare
Phrase Grounding Visual Genome GbS VG Pointing Game Accuracy 55.91 Detector-Free Weakly Supervised Grounding by Separation aarbelle/GroundingBySeparation 3 Compare
Image Generation from Scene Graphs Visual Genome 64x64 MIGS FID 54.24 MIGS: Meta Image Generation from Scene Graphs migs2021/migs 2 Compare
Unsupervised KG-to-Text Generation VG graph-text GT-BT (composed noise) BLEU 23.2 An Unsupervised Joint System for Text Generation from... mnschmit/unsupervised-graph-text-conversion 1 Compare
Unsupervised semantic parsing VG graph-text GT-BT (composed noise) F1 21.7 An Unsupervised Joint System for Text Generation from... mnschmit/unsupervised-graph-text-conversion 1 Compare
Visual Question Answering (VQA) Visual Genome (subjects) CMN Percentage correct 44.24 Modeling Relationships in Referential Expressions with... hengyuan-hu/bottom-up-attention-vqa +1 1 Compare
Visual Question Answering (VQA) Visual Genome (pairs) CMN Percentage correct 28.52 Modeling Relationships in Referential Expressions with... hengyuan-hu/bottom-up-attention-vqa +1 1 Compare
Visual Relationship Detection Visual Genome PEVL R@100 66.3 PEVL: Position-enhanced Pre-training and Prompt Tuning... thunlp/pevl 1 Compare

Papers archive 2025-07-28

30 shown of 44 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 1,256. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
KnowZRel: Common Sense Knowledge-based Zero-Shot Relationship Retrieval for Generalised Scene Graph Generation 1 2 21 Feb 2025 not harvested
Semantic Diversity-aware Prototype-based Learning for Unbiased Scene Graph Generation 1 3 22 Jul 2024 ran 1 of 1 samples (0 unverified; 1 pointer-only for licence)
Groupwise Query Specialization and Quality-Aware Multi-Assignment for Transformer-based Visual Relationship Detection 1 2 26 Mar 2024 ran 1 of 3 samples (2 unverified; 3 pointer-only for licence)
ControlCap: Controllable Region-level Captioning 1 1 31 Jan 2024 ran 3 of 4 samples (1 unverified; 4 pointer-only for licence)
NeuSyRE: Neuro-Symbolic Visual Understanding and Reasoning Framework based on Scene Graph Enrichment 1 1 5 Nov 2023 not harvested
Panoptic Scene Graph Generation with Semantics-Prototype Learning 1 1 28 Jul 2023 not harvested
LayoutDiffusion: Controllable Diffusion Model for Layout-to-image Generation 2 2 30 Mar 2023 ran 5 of 19 samples (14 unverified)
GRiT: A Generative Region-to-text Transformer for Object Understanding 1 1 1 Dec 2022 ran 5 of 8 samples (3 unverified; 8 pointer-only for licence)
An Enhanced Object Detection Model for Scene Graph Generation 0 1 18 Nov 2022 not harvested
Expressive Scene Graph Generation Using Commonsense Knowledge Infusion for Visual Understanding and Reasoning 1 1 31 May 2022 not harvested
Dual-Perspective Semantic-Aware Representation Blending for Multi-Label Image Recognition with Partial Labels 1 1 26 May 2022 not harvested
PEVL: Position-enhanced Pre-training and Prompt Tuning for Vision-language Models 1 1 23 May 2022 ran 3 of 7 samples (4 unverified)
Heterogeneous Semantic Transfer for Multi-label Recognition with Partial Labels 1 1 23 May 2022 not harvested
Fine-Grained Scene Graph Generation with Data Transfer 2 4 22 Mar 2022 not harvested
Stacked Hybrid-Attention and Group Collaborative Learning for Unbiased Scene Graph Generation 1 1 18 Mar 2022 ran 3 of 7 samples (4 unverified; 7 pointer-only for licence)
Biasing Like Human: A Cognitive Bias Framework for Scene Graph Generation 1 1 17 Mar 2022 not harvested
Semantic-Aware Representation Blending for Multi-Label Image Recognition with Partial Labels 1 1 4 Mar 2022 ran 6 of 7 samples (1 unverified; 7 pointer-only for licence)
Interactive Image Synthesis with Panoptic Layout Generation 1 2 4 Mar 2022 ran 0 of 11 samples (11 unverified)
Structured Semantic Transfer for Multi-Label Recognition with Partial Labels 1 1 21 Dec 2021 not harvested
MIGS: Meta Image Generation from Scene Graphs 1 2 22 Oct 2021 not harvested
Recovering the Unbiased Scene Graphs from the Biased Ones 1 7 5 Jul 2021 not harvested
Tackling the Challenges in Scene Graph Generation with Local-to-Global Interactions 1 1 16 Jun 2021 not harvested
Detector-Free Weakly Supervised Grounding by Separation 1 2 20 Apr 2021 not harvested
Energy-Based Learning for Scene Graph Generation 1 1 3 Mar 2021 not harvested
CogTree: Cognition Tree Loss for Unbiased Scene Graph Generation 1 7 16 Sep 2020 not harvested
PCPL: Predicate-Correlation Perception Learning for Unbiased Scene Graph Generation 1 6 2 Sep 2020 not harvested
GPS-Net: Graph Property Sensing Network for Scene Graph Generation 1 1 29 Mar 2020 ran 2 of 8 samples (6 unverified)
Learning Layout and Style Reconfigurable GANs for Controllable Image Synthesis 3 1 25 Mar 2020 not harvested
Object-Centric Image Generation from Layouts 0 3 16 Mar 2020 not harvested
Unbiased Scene Graph Generation from Biased Training 6 7 27 Feb 2020 ran 3 of 6 samples (3 unverified; 6 pointer-only for licence)

The full list of 44 is in the JSON twin.

Dataset loaders archive 2025-07-28

2 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

CC BY 4.0

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • VG graph-text
  • Visual Genome 256x256
  • Visual Genome 64x64
  • Visual Genome 128x128
  • Visual Genome (subjects)
  • Visual Genome (pairs)
  • Visual Genome

7 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections