Browse State-of-the-Art › Spatial Reasoning › Papers, page 3
Spatial Reasoning
Papers archive 2025-07-28
archive papers tagged: 453 · with a code link: 198 · where Syntology ran a sample: 70 (59 with a run with no instrument failure, 11 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (70 of 453 tagged: 59 with a run with no instrument failure, 11 where every run was a failure of Syntology's instrument)
Page 3 of 5: papers 201 to 300 of 453, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
ByDeWay: Boost Your multimodal LLM with DEpth prompting in a Training-Free Way11 Jul 2025 0 repositories listed
-
M2-Reasoning: Empowering MLLMs with Unified General and Spatial Reasoning11 Jul 2025 0 repositories listed
-
A Neural Representation Framework with LLM-Driven Spatial Reasoning for Open-Vocabulary 3D Visual Grounding9 Jul 2025 0 repositories listed
-
Optimising Language Models for Downstream Tasks: A Post-Training Perspective26 Jun 2025 0 repositories listed
-
ShotBench: Expert-Level Cinematic Understanding in Vision-Language Models26 Jun 2025 0 repositories listed
-
World-aware Planning Narratives Enhance Large Vision-Language Model Planner26 Jun 2025 0 repositories listed
-
From 2D to 3D Cognition: A Brief Survey of General World Models25 Jun 2025 0 repositories listed
-
Video Perception Models for 3D Scene Synthesis25 Jun 2025 0 repositories listed
-
ImmerseGen: Agent-Guided Immersive World Generation with Alpha-Textured Proxies17 Jun 2025 0 repositories listed
-
PeRL: Permutation-Enhanced Reinforcement Learning for Interleaved Vision-Language Reasoning17 Jun 2025 0 repositories listed
-
SIRI-Bench: Challenging VLMs' Spatial Intelligence through Complex Reasoning Tasks17 Jun 2025 0 repositories listed
-
Leveraging LLMs for Mission Planning in Precision Agriculture11 Jun 2025 0 repositories listed
-
A Multi-Modal Spatial Risk Framework for EV Charging Infrastructure Using Remote Sensing10 Jun 2025 0 repositories listed
-
PhyBlock: A Progressive Benchmark for Physical Understanding and Planning via 3D Block Assembly10 Jun 2025 0 repositories listed
-
Direct Numerical Layout Generation for 3D Indoor Scene Synthesis via Spatial Reasoning5 Jun 2025 0 repositories listed
-
From Objects to Anywhere: A Holistic Benchmark for Multi-level Visual Grounding in 3D Scenes5 Jun 2025 0 repositories listed
-
RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics4 Jun 2025 0 repositories listed
-
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing4 Jun 2025 0 repositories listed
-
OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models3 Jun 2025 0 repositories listed
-
ReSpace: Text-Driven 3D Scene Synthesis and Editing with Preference Alignment3 Jun 2025 0 repositories listed
-
Out of Sight, Not Out of Context? Egocentric Spatial Reasoning in VLMs Across Disjoint Frames30 May 2025 0 repositories listed
-
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces30 May 2025 0 repositories listed
-
MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence29 May 2025 0 repositories listed
-
Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence29 May 2025 0 repositories listed
-
Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models27 May 2025 0 repositories listed
-
VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models27 May 2025 0 repositories listed
-
Agentic 3D Scene Generation with Spatially Contextualized VLMs26 May 2025 0 repositories listed
-
MEBench: A Novel Benchmark for Understanding Mutual Exclusivity Bias in Vision-Language Models26 May 2025 0 repositories listed
-
ViTaPEs: Visuotactile Position Encodings for Cross-Modal Alignment in Multimodal Transformers26 May 2025 0 repositories listed
-
Can MLLMs Guide Me Home? A Benchmark Study on Fine-Grained Visual Reasoning from Transit Maps24 May 2025 0 repositories listed
-
Towards Dynamic 3D Reconstruction of Hand-Instrument Interaction in Ophthalmic Surgery23 May 2025 0 repositories listed
-
23 May 2025 0 repositories listed
-
MEgoHand: Multimodal Egocentric Hand-Object Interaction Motion Generation22 May 2025 0 repositories listed
-
MMMR: Benchmarking Massive Multi-Modal Reasoning Tasks22 May 2025 0 repositories listed
-
SEM: Enhancing Spatial Understanding for Robust Robot Manipulation22 May 2025 0 repositories listed
-
VLM-R³: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought22 May 2025 0 repositories listed
-
ReGUIDE: Data Efficient GUI Grounding via Spatial Reasoning and Search21 May 2025 0 repositories listed
-
STAR-R1: Spacial TrAnsformation Reasoning by Reinforcing Multimodal LLMs21 May 2025 0 repositories listed
-
From Templates to Natural Language: Generalization Challenges in Instruction-Tuned LLMs for Spatial Reasoning20 May 2025 0 repositories listed
-
Towards Embodied Cognition in Robots via Spatially Grounded Synthetic Worlds20 May 2025 0 repositories listed
-
Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation19 May 2025 0 repositories listed
-
Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind18 May 2025 0 repositories listed
-
SSR: Enhancing Depth Perception in Vision-Language Models via Rationale-Guided Spatial Reasoning18 May 2025 0 repositories listed
-
Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning?17 May 2025 0 repositories listed
-
PRS-Med: Position Reasoning Segmentation with Vision-Language Model in Medical Imaging17 May 2025 0 repositories listed
-
A Light and Smart Wearable Platform with Multimodal Foundation Model for Enhanced Spatial Reasoning in People with Blindness and Low Vision16 May 2025 0 repositories listed
-
SITE: towards Spatial Intelligence Thorough Evaluation8 May 2025 0 repositories listed
-
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models8 May 2025 0 repositories listed
-
Preliminary Explorations with GPT-4o(mni) Native Image Generation6 May 2025 0 repositories listed
-
Beyond Recognition: Evaluating Visual Perspective Taking in Vision Language Models3 May 2025 0 repositories listed
-
FreeInsert: Disentangled Text-Guided Object Insertion in 3D Gaussian Scene without Spatial Priors2 May 2025 0 repositories listed
-
SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models1 May 2025 0 repositories listed
-
First Order Logic with Fuzzy Semantics for Describing and Recognizing Nerves in Medical Images30 Apr 2025 0 repositories listed
-
SpatialReasoner: Towards Explicit and Generalizable 3D Spatial Reasoning28 Apr 2025 0 repositories listed
-
A Review of 3D Object Detection with Vision-Language Models25 Apr 2025 0 repositories listed
-
Spatial Reasoner: A 3D Inference Pipeline for XR Applications25 Apr 2025 0 repositories listed
-
A Call for New Recipes to Enhance Spatial Reasoning in MLLMs21 Apr 2025 0 repositories listed
-
EarthGPT-X: Enabling MLLMs to Flexibly and Comprehensively Understand Multi-Source Remote Sensing Imagery17 Apr 2025 0 repositories listed
-
Intelligence of Things: A Spatial Context-Aware Control System for Smart Devices16 Apr 2025 0 repositories listed
-
Embodied World Models Emerge from Navigational Task in Open-Ended Environments15 Apr 2025 0 repositories listed
-
LVLM_CSP: Accelerating Large Vision Language Models via Clustering, Scattering, and Pruning for Reasoning Segmentation15 Apr 2025 0 repositories listed
-
A Survey of Large Language Model-Powered Spatial Intelligence Across Scales: Advances in Embodied Agents, Smart Cities, and Earth Science14 Apr 2025 0 repositories listed
-
Foundation Models for Remote Sensing: An Analysis of MLLMs for Object Localization14 Apr 2025 0 repositories listed
-
Perturbed State Space Feature Encoders for Optical Flow with Event Cameras14 Apr 2025 0 repositories listed
-
VisualPuzzles: Decoupling Multimodal Reasoning Evaluation from Domain Knowledge14 Apr 2025 0 repositories listed
-
Embodied Chain of Action Reasoning with Multi-Modal Foundation Model for Humanoid Loco-manipulation13 Apr 2025 0 repositories listed
-
VisuoThink: Empowering LVLM Reasoning with Multimodal Tree Search12 Apr 2025 0 repositories listed
-
AerialVG: A Challenging Benchmark for Aerial Visual Grounding by Exploring Positional Relations10 Apr 2025 0 repositories listed
-
Endowing Embodied Agents with Spatial Reasoning Capabilities for Vision-and-Language Navigation9 Apr 2025 0 repositories listed
-
How to Enable LLM with 3D Capacity? A Survey of Spatial Reasoning in LLM8 Apr 2025 0 repositories listed
-
Towards Visual Text Grounding of Multimodal Large Language Model7 Apr 2025 0 repositories listed
-
Advancing Egocentric Video Question Answering with Multimodal Large Language Models6 Apr 2025 0 repositories listed
-
NuScenes-SpatialQA: A Spatial Understanding and Reasoning Benchmark for Vision-Language Models in Autonomous Driving4 Apr 2025 0 repositories listed
-
RSRWKV: A Linear-Complexity 2D Attention Mechanism for Efficient Remote Sensing Vision Task26 Mar 2025 0 repositories listed
-
DataPlatter: Boosting Robotic Manipulation Generalization with Minimal Costly Data25 Mar 2025 0 repositories listed
-
LEGO-Puzzles: How Good Are MLLMs at Multi-Step Spatial Reasoning?25 Mar 2025 0 repositories listed
-
ST-VLM: Kinematic Instruction Tuning for Spatio-Temporal Reasoning in Vision-Language Models25 Mar 2025 0 repositories listed
-
Aether: Geometric-Aware Unified World Modeling24 Mar 2025 0 repositories listed
-
AlphaSpace: Enabling Robotic Actions through Semantic Tokenization and Symbolic Reasoning24 Mar 2025 0 repositories listed
-
MLLM-For3D: Adapting Multimodal Large Language Model for 3D Reasoning Segmentation23 Mar 2025 0 repositories listed
-
Beyond Semantics: Rediscovering Spatial Awareness in Vision-Language Models21 Mar 2025 0 repositories listed
-
A Vision Centric Remote Sensing Benchmark20 Mar 2025 0 repositories listed
-
OmniGeo: Towards a Multimodal Large Language Models for Geospatial Artificial Intelligence20 Mar 2025 0 repositories listed
-
Statistical applications of the 20/60/20 rule in risk management and portfolio optimization19 Mar 2025 0 repositories listed
-
UI-Vision: A Desktop-centric GUI Benchmark for Visual Perception and Interaction19 Mar 2025 0 repositories listed
-
EmbodiedVSR: Dynamic Scene Graph-Guided Chain-of-Thought Reasoning for Visual Spatial Tasks14 Mar 2025 0 repositories listed
-
CleverDistiller: Simple and Spatially Consistent Cross-modal Distillation12 Mar 2025 0 repositories listed
-
Boosting Diffusion-Based Text Image Super-Resolution Model Towards Generalized Real-World Scenarios10 Mar 2025 0 repositories listed
-
Navigating Motion Agents in Dynamic and Cluttered Environments through LLM Reasoning10 Mar 2025 0 repositories listed
-
An Empirical Study of Conformal Prediction in LLM with ASP Scaffolds for Robust Reasoning7 Mar 2025 0 repositories listed
-
ROCKET-2: Steering Visuomotor Policy via Cross-View Goal Alignment4 Mar 2025 0 repositories listed
-
VisFactor: Benchmarking Fundamental Visual Cognition in Multimodal Large Language Models23 Feb 2025 0 repositories listed
-
Mem2Ego: Empowering Vision-Language Models with Global-to-Ego Memory for Long-Horizon Embodied Navigation20 Feb 2025 0 repositories listed
-
Large Language Models and Mathematical Reasoning Failures17 Feb 2025 0 repositories listed
-
Large Language-Geometry Model: When LLM meets Equivariance16 Feb 2025 0 repositories listed
-
STMA: A Spatio-Temporal Memory Agent for Long-Horizon Embodied Task Planning14 Feb 2025 0 repositories listed
-
A Solver-Aided Hierarchical Language for LLM-Driven CAD Design13 Feb 2025 0 repositories listed
-
Visual Agentic AI for Spatial Reasoning with a Dynamic API10 Feb 2025 0 repositories listed
-
Vision-Integrated LLMs for Autonomous Driving Assistance : Human Performance Comparison and Trust Evaluation6 Feb 2025 0 repositories listed
-
A Schema-Guided Reason-while-Retrieve framework for Reasoning on Scene Graphs with Large-Language-Models (LLMs)5 Feb 2025 0 repositories listed