Papers › SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

22 Jan 2024CVPR 2024 1arXiv:2401.12168archive 2025-07-28

Boyuan Chen, Zhuo Xu, Sean Kirmani, Brian Ichter, Danny Driess, Pete Florence, Dorsa Sadigh, Leonidas Guibas, Fei Xia

Understanding and reasoning about spatial relationships is a fundamental capability for Visual Question Answering (VQA) and robotics. While Vision Language Models (VLM) have demonstrated remarkable performance in certain VQA benchmarks, they still lack capabilities in 3D spatial reasoning, such as recognizing quantitative relationships of physical objects like distances or size differences. We hypothesize that VLMs' limited spatial reasoning capability is due to the lack of 3D spatial knowledge in training data and aim to solve this problem by training VLMs with Internet-scale spatial reasoning data. To this end, we present a system to facilitate this approach. We first develop an automatic 3D spatial VQA data generation framework that scales up to 2 billion VQA examples on 10 million real-world images. We then investigate various factors in the training recipe, including data quality, training pipeline, and VLM architecture. Our work features the first internet-scale 3D spatial reasoning dataset in metric space. By training a VLM on such data, we significantly enhance its ability on both qualitative and quantitative spatial VQA. Finally, we demonstrate that this VLM unlocks novel downstream applications in chain-of-thought spatial reasoning and robotics due to its quantitative estimation capability. Project website: https://spatial-vlm.github.io/

PaperPDFConference PDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Question AnsweringSpatial ReasoningVisual Question AnsweringVisual Question Answering (VQA)

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Spatial Reasoning 6-DoF SpatialBench SpaceMantis Orientation-abs 25.0 #5 of 7 Archive leaderboard report
Spatial Reasoning 6-DoF SpatialBench SpaceMantis Orientation-rel 27.2 #5 of 7 Archive leaderboard report
Spatial Reasoning 6-DoF SpatialBench SpaceMantis Position-abs 29.2 #5 of 7 Archive leaderboard report
Spatial Reasoning 6-DoF SpatialBench SpaceMantis Position-rel 33.6 #5 of 7 Archive leaderboard report
Spatial Reasoning 6-DoF SpatialBench SpaceMantis Total 28.9 #5 of 7 Archive leaderboard report
Spatial Reasoning 6-DoF SpatialBench SpaceLLaVA Orientation-abs 24.9 #6 of 7 Archive leaderboard report
Spatial Reasoning 6-DoF SpatialBench SpaceLLaVA Orientation-rel 30.9 #6 of 7 Archive leaderboard report
Spatial Reasoning 6-DoF SpatialBench SpaceLLaVA Position-abs 30.5 #6 of 7 Archive leaderboard report
Spatial Reasoning 6-DoF SpatialBench SpaceLLaVA Position-rel 32.4 #6 of 7 Archive leaderboard report
Spatial Reasoning 6-DoF SpatialBench SpaceLLaVA Total 28.2 #6 of 7 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections