Datasets › SQA3D

SQA3D (Situated Question Answering in 3D Scenes)

Introduced by Xiaojian Ma et al. in SQA3D: Situated Question Answering in 3D Scenes30 Jan 2023 archive 2025-07-28

SQA3D is a dataset for embodied scene understanding, where an agent needs to comprehend the scene it situates from an first person's perspective and answer questions. The questions are designed to be situated, embodied and knowledge-intensive. We offer three different modalities to represent a 3D scene: 3D scan, egocentric video and BEV picture.

Benchmarks archive 2025-07-28

All 3 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
3D Question Answering (3D-QA) SQA3D LLaVA-3D Exact Match 60.1 LLaVA-3D: A Simple yet Effective Pathway to Empowering... — 13 Compare
Question Answering SQA3D CREMA AnswerExactMatch (Question Answering) 54.6 CREMA: Generalizable and Efficient Video-Language... Yui010206/CREMA 7 Compare
Referring Expression SQA3D Random Acc@0.5m 14.60 SQA3D: Situated Question Answering in 3D Scenes SilongYong/SQA3D 1 Compare

Papers archive 2025-07-28

18 shown of 18 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 58. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding 1 1 30 Nov 2024 not harvested
Video Instruction Tuning With Synthetic Data 0 1 3 Oct 2024 not harvested
LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness 0 1 26 Sep 2024 not harvested
Lexicon3D: Probing Visual Foundation Models for Complex 3D Scene Understanding 1 1 5 Sep 2024 ran 2 of 9 samples (7 unverified)
LLaVA-OneVision: Easy Visual Task Transfer 2 1 6 Aug 2024 not harvested
Situational Awareness Matters in 3D Vision Language Reasoning 1 1 11 Jun 2024 ran 10 of 10 samples (0 unverified)
Unifying 3D Vision-Language Understanding via Promptable Queries 0 1 19 May 2024 not harvested
Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning 0 1 18 Mar 2024 not harvested
CREMA: Generalizable and Efficient Video-Language Reasoning via Multimodal Modular Fusion 1 1 8 Feb 2024 ran 8 of 10 samples (2 unverified)
Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers 2 2 13 Dec 2023 ran 11 of 11 samples (0 unverified; 5 pointer-only for licence)
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark 3 1 28 Nov 2023 ran 7 of 10 samples (3 unverified)
An Embodied Generalist Agent in 3D World 1 1 18 Nov 2023 ran 3 of 13 samples (10 unverified)
Frozen Transformers in Language Models Are Effective Visual Encoder Layers 2 1 19 Oct 2023 ran 8 of 16 samples (8 unverified)
3D-VisTA: Pre-trained Transformer for 3D Vision and Text Alignment 1 1 8 Aug 2023 ran 4 of 6 samples (2 unverified)
SQA3D: Situated Question Answering in 3D Scenes 1 3 14 Oct 2022 ran 6 of 8 samples (2 unverified)
ScanQA: 3D Question Answering for Spatial Scene Understanding 1 1 20 Dec 2021 not harvested
Scan2Cap: Context-aware Dense Captioning in RGB-D Scans 0 1 3 Dec 2020 not harvested
Deep Modular Co-Attention Networks for Visual Question Answering 7 1 25 Jun 2019 ran 0 of 1 samples (1 unverified)

Dataset loaders archive 2025-07-28

1 loader as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

CC-BY-4.0

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • SQA3D

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections