Browse State-of-the-Art › Spatial Reasoning
Spatial Reasoning
198 papers with code · 2 benchmarks · 5 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
16 leaderboard tables shown for this task, 2 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 16 until expanded.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| 6-DoF SpatialBench (7 rows) | SoFar | SoFar: Language-Grounded Orientation Bridges Spatial Reasoning and... | code | — | Compare |
| EmbSpatial-Bench (5 rows) | SoFar | SoFar: Language-Grounded Orientation Bridges Spatial Reasoning and... | code | — | Compare |
| 3DSRBench (0 rows) | no rows in the archive | — | — | ||
| BlINK (0 rows) | no rows in the archive | — | — | ||
| cvbench (0 rows) | no rows in the archive | — | — | ||
| MMIU (0 rows) | no rows in the archive | — | — | ||
| MMVP (0 rows) | no rows in the archive | — | — | ||
| Q-Spatial-Bench (0 rows) | no rows in the archive | — | — | ||
| QSpatialBench-Plus (0 rows) | no rows in the archive | — | — | ||
| QSpatialBench-ScanNet (0 rows) | no rows in the archive | — | — | ||
| RealWorldQA (0 rows) | no rows in the archive | — | — | ||
| SpatialBench (0 rows) | no rows in the archive | — | — | ||
| SpatialSense (0 rows) | no rows in the archive | — | — | ||
| VGBench (0 rows) | no rows in the archive | — | — | ||
| VSI-Bench_8 (0 rows) | no rows in the archive | — | — | ||
| VSR-ZeroShot (0 rows) | no rows in the archive | — | — | ||
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
5 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 198 papers with code (453 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
13 Apr 2017 35 repositories listedOn the other hand, modeling object-object relationships requires {\bf spatial} reasoning -- not only do we need a memory to store the spatial layout, but also a effective reasoning module to extract spatial patterns.
-
17 Apr 2023 13 repositories listed Syntology ran 16 of 51 samples · 35 unverifiedInstruction tuning large language models (LLMs) using machine-generated instruction-following data has improved zero-shot capabilities on new tasks, but the idea is less explored in the multimodal field.
-
15 Mar 2023 11 repositories listed Syntology ran 2 of 5 samples · 3 unverified · 1 pointer-only (licence)We report the development of GPT-4, a large-scale, multimodal model which can accept image and text inputs and produce text outputs.
-
5 Oct 2023 9 repositories listed Syntology ran 6 of 9 samples · 3 unverified · 8 pointer-only (licence)Large multimodal models (LMM) have recently shown encouraging progress with visual instruction tuning.
-
20 Apr 2023 6 repositories listedOur work, for the first time, uncovers that properly aligning the visual features with an advanced large language model can possess numerous advanced multi-modal abilities demonstrated by GPT-4, such as detailed image…
-
8 Nov 2020 5 repositories listed Syntology ran 2 of 3 samples · 1 unverifiedIn the recent months, a wide spectrum of efficient, fast Transformers have been proposed to tackle this problem, more often than not claiming superior or comparable model quality to vanilla Transformer models.
-
30 Apr 2022 4 repositories listed Syntology ran 5 of 10 samples · 5 unverifiedSpatial relations are a basic part of human cognition.
-
29 Nov 2018 4 repositories listedWe study the problem of jointly reasoning about language and vision through a navigation and spatial reasoning task.
-
23 Nov 2016 4 repositories listed Syntology ran 0 of 3 samples · 3 unverifiedOur key contribution is the collection of a large-scale dataset consisting of 150K human-played games with a total of 800K visual question-answer pairs on 66K images.
-
31 Dec 2024 3 repositories listed Syntology ran 3 of 3 samples · 0 unverified · 3 pointer-only (licence)To bridge this gap, we introduce MapEval, a benchmark designed to assess diverse and complex map-based user queries with geo-spatial reasoning.
-
25 Nov 2024 3 repositories listed Syntology ran 0 of 9 samples · 9 unverifiedRecent advancements in artificial intelligence have sparked interest in scientific assistants that could support researchers across the full spectrum of scientific workflows, from literature review to experimental…
-
23 Feb 2018 3 repositories listedIn this paper, we present a modular framework for tracking multiple objects (vehicles), capable of accepting object proposals from different sensor modalities (vision and range) and a variable number of sensors, to…
-
19 Apr 2025 2 repositories listed Syntology ran 0 of 3 samples · 3 unverifiedRecent works have begun exploring reasoning in GUI tasks with encouraging results.
-
2 Apr 2025 2 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)Motivated by the success of Reinforcement Learning with Verifiable Reward (RLVR) in unlocking LLM reasoning abilities, this work aims to improve MLLMs in video spatial reasoning through the RLVR paradigm.
-
20 Mar 2025 2 repositories listed Syntology ran 5 of 19 samples · 14 unverifiedIn this paper, we question whether we have a reliable self-supervised point cloud model that can be used for diverse 3D tasks via simple linear probing, even with limited data and minimal computation.
-
20 Mar 2025 2 repositories listedWith this benchmark, we aim to provide a resource for 3D scene understanding that aids the development of robust, interactive navigation systems.
-
18 Feb 2025 2 repositories listedSpatial intelligence is a critical component of embodied AI, promoting robots to understand and interact with their environments.
-
18 Feb 2025 2 repositories listed Syntology ran 2 of 4 samples · 2 unverified · 4 pointer-only (licence)To bridge this gap, we introduce CityEQA, a new task where an embodied agent answers open-vocabulary questions through active exploration in dynamic city spaces.
-
30 Sep 2024 2 repositories listedRecent advancements in Large Language Models (LLMs) have showcased their ability to perform complex reasoning tasks, but their effectiveness in planning remains underexplored.
-
23 May 2024 2 repositories listedSpatial reasoning plays a vital role in both human cognition and machine intelligence, prompting new research into language models' (LMs) capabilities in this regard.
-
12 Jan 2024 2 repositories listedIn this work we propose a road segmentation benchmark dataset, Chesapeake Roads Spatial Context (RSC), for evaluating the spatial long-range context understanding of geospatial machine learning models and show how…
-
30 Oct 2023 2 repositories listed Syntology ran 5 of 9 samples · 4 unverified · 2 pointer-only (licence)Recent vision-language (VL) models are powerful, but can they reliably distinguish "right" from "left"?
-
Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond24 Aug 2023 2 repositories listed Syntology ran 0 of 2 samples · 2 unverified · 2 pointer-only (licence)In this work, we introduce the Qwen-VL series, a set of large-scale vision-language models (LVLMs) designed to perceive and understand both texts and images.
-
17 Aug 2023 2 repositories listedThis paper presents Chat-3D, which combines the 3D visual perceptual ability of pre-trained 3D representations and the impressive reasoning and conversation capabilities of advanced LLMs to achieve the first universal…
-
30 Jun 2023 2 repositories listed3D perceptual representations are well suited for robot manipulation as they easily encode occlusions and simplify spatial reasoning.
-
23 May 2023 2 repositories listedOur method significantly outperforms the base diffusion model and several strong baselines in accurately generating images according to prompts that require various capabilities, doubling the generation accuracy across…
-
12 Apr 2022 2 repositories listed Syntology ran 0 of 5 samples · 5 unverifiedTraining a referring expression comprehension (ReC) model for a new visual domain requires collecting referring expressions, and potentially corresponding bounding boxes, for images in the domain.
-
13 Jul 2021 2 repositories listed Syntology ran 1 of 7 samples · 6 unverifiedIn the context of visual navigation, the capacity to map a novel environment is necessary for an agent to exploit its observation history in the considered place and efficiently reach known goals.
-
1 Jun 2021 2 repositories listedThis paper proposes a question-answering (QA) benchmark for spatial reasoning on natural language text which contains more realistic spatial phenomena not covered by prior work and is challenging for state-of-the-art…
-
15 Feb 2021 2 repositories listedSpatial memory, or the ability to remember and recall specific locations and objects, is central to autonomous agents' ability to carry out tasks in real environments.
Syntology lines on 16 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections