Datasets › GQA

GQA

Introduced by Drew A. Hudson et al. in GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering1 Jan 2019 archive 2025-07-28

The GQA dataset is a large-scale visual question answering dataset with real images from the Visual Genome dataset and balanced question-answer pairs. Each training and validation image is also associated with scene graph annotations describing the classes and attributes of those objects in the scene, and their pairwise relations. Along with the images and question-answer pairs, the GQA dataset provides two types of pre-extracted visual features for each image – convolutional grid features of size 7×7×2048 extracted from a ResNet-101 network trained on ImageNet, and object detection features of size Ndet×2048 (where Ndet is the number of detected objects in each image with a maximum of 100 per image) from a Faster R-CNN detector.

Source: Language-Conditioned Graph Networks for Relational Reasoning Image Source: https://arxiv.org/pdf/1902.09506.pdf

Benchmarks archive 2025-07-28

All 8 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

Papers archive 2025-07-28

24 shown of 24 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 749. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
KnowZRel: Common Sense Knowledge-based Zero-Shot Relationship Retrieval for Generalised Scene Graph Generation 1 2 21 Feb 2025 not harvested
CuMo: Scaling Multimodal LLM with Co-Upcycled Mixture-of-Experts 1 1 9 May 2024 ran 10 of 12 samples (2 unverified)
Learning to Localize Objects Improves Spatial Reasoning in Visual-LLMs 1 1 11 Apr 2024 not harvested
HYDRA: A Hyper Agent for Dynamic Compositional Visual Reasoning 1 1 19 Mar 2024 ran 5 of 8 samples (3 unverified)
Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization 1 1 5 Feb 2024 ran 3 of 5 samples (2 unverified; 5 pointer-only for licence)
Lyrics: Boosting Fine-grained Language-Vision Alignment and Comprehension via Semantic-aware Visual Objects 0 1 8 Dec 2023 not harvested
Visual Program Distillation: Distilling Tools and Programmatic Reasoning into Vision-Language Models 0 1 5 Dec 2023 not harvested
VinVL+L: Enriching Visual Representation with Location Context in VQA 1 1 22 Feb 2023 not harvested
BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models 17 6 30 Jan 2023 ran 4 of 8 samples (4 unverified; 1 pointer-only for licence)
Plug-and-Play VQA: Zero-shot VQA by Conjoining Large Pretrained Models with Zero Training 3 1 17 Oct 2022 ran 2 of 3 samples (1 unverified; 3 pointer-only for licence)
PEVL: Position-enhanced Pre-training and Prompt Tuning for Vision-language Models 1 1 23 May 2022 ran 3 of 7 samples (4 unverified)
RelViT: Concept-guided Vision Transformer for Visual Relational Reasoning 1 1 24 Apr 2022 ran 6 of 10 samples (4 unverified; 10 pointer-only for licence)
A Good Prompt Is Worth Millions of Parameters: Low-resource Prompt-based Learning for Vision-Language Models 1 1 16 Oct 2021 not harvested
Coarse-to-Fine Reasoning for Visual Question Answering 2 1 6 Oct 2021 ran 4 of 9 samples (5 unverified)
ProTo: Program-Guided Transformer for Program-Guided Tasks 1 1 2 Oct 2021 ran 0 of 1 samples (1 unverified)
MDETR -- Modulated Detection for End-to-End Multi-Modal Understanding 5 1 26 Apr 2021 ran 6 of 11 samples (5 unverified)
GraghVQA: Language-Guided Graph Neural Networks for Graph-based Visual Question Answering 1 1 20 Apr 2021 not harvested
VinVL: Revisiting Visual Representations in Vision-Language Models 7 1 2 Jan 2021 ran 2 of 2 samples (0 unverified; 1 pointer-only for licence)
LXMERT: Learning Cross-Modality Encoder Representations from Transformers 9 4 20 Aug 2019 ran 4 of 15 samples (11 unverified; 3 pointer-only for licence)
Bilinear Graph Networks for Visual Question Answering 0 1 23 Jul 2019 not harvested
Learning by Abstraction: The Neural State Machine 4 2 9 Jul 2019 ran 3 of 21 samples (18 unverified; 2 pointer-only for licence)
Language-Conditioned Graph Networks for Relational Reasoning 1 2 10 May 2019 ran 0 of 10 samples (10 unverified)
GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering 5 2 25 Feb 2019 not harvested
Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering 65 1 25 Jul 2017 ran 9 of 9 samples (0 unverified; 6 pointer-only for licence)

Dataset loaders archive 2025-07-28

1 loader as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

CC BY 4.0

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • GQA test-std
  • GQA test-dev
  • GQA Test2019
  • GQA
  • GQA-OOD

5 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections