Datasets › ScanQA

ScanQA (ScanQA: 3D Question Answering for Spatial Scene Understanding)

Introduced by Daichi Azuma et al. in ScanQA: 3D Question Answering for Spatial Scene Understanding20 Dec 2021 archive 2025-07-28

We collected 41,363 questions and 58,191 answers, in- cluding 32,337 unique questions and 16,999 unique an- swers. Table 2 presents the statistics of the ScanQA dataset. This dataset is an order of magnitude larger than existing embodied question-answering datasets in terms of both question size and variation. For example, the EQA dataset contains 4,246 questions, consisting of 147 unique questions in its training set. The EQA-MP3D dataset contains 767 questions consisting of 174 unique questions in its training set. Considering that our dataset contains not only question–answer pairs but also 3D object localization annotations, we assume that this is the largest dataset to specify the nature of objects in 3D scenes with the question answering form. The distribution of the questions based on their first word. We collected various types of questions through question auto-generation and editing by humans.

Benchmarks archive 2025-07-28

All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
3D Question Answering (3D-QA) ScanQA Test w/ objects BridgeQA Exact Match 31.29 Bridging the Gap between 2D and 3D Visual Question... matthewdm0816/bridgeqa 18 Compare

Papers archive 2025-07-28

13 shown of 13 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 70. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding 1 1 30 Nov 2024 not harvested
LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness 0 1 26 Sep 2024 not harvested
LLaVA-OneVision: Easy Visual Task Transfer 2 1 6 Aug 2024 not harvested
Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning 0 1 18 Mar 2024 not harvested
Bridging the Gap between 2D and 3D Visual Question Answering: A Fusion Approach for 3D VQA 1 1 24 Feb 2024 ran 10 of 13 samples (3 unverified; 13 pointer-only for licence)
Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers 2 2 13 Dec 2023 ran 11 of 11 samples (0 unverified; 5 pointer-only for licence)
Towards Learning a Generalist Model for Embodied Navigation 2 1 4 Dec 2023 ran 10 of 15 samples (5 unverified)
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark 3 1 28 Nov 2023 ran 7 of 10 samples (3 unverified)
An Embodied Generalist Agent in 3D World 1 1 18 Nov 2023 ran 3 of 13 samples (10 unverified)
3D-VisTA: Pre-trained Transformer for 3D Vision and Text Alignment 1 1 8 Aug 2023 ran 4 of 6 samples (2 unverified)
3D-LLM: Injecting the 3D World into Large Language Models 5 3 24 Jul 2023 ran 5 of 10 samples (5 unverified; 10 pointer-only for licence)
Visual Instruction Tuning 13 1 17 Apr 2023 ran 16 of 51 samples (35 unverified)
ScanQA: 3D Question Answering for Spatial Scene Understanding 1 3 20 Dec 2021 not harvested

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

No modality tagged.

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • ScanQA Test w/ objects
  • ScanQA

2 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections