{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/why-is-spatial-reasoning-hard-for-vlms-an","title":"Why Is Spatial Reasoning Hard for VLMs? An Attention Mechanism Perspective on Focus Areas","arxiv_id":"2503.01773","date":"2025-03-03","proceeding":null,"authors":["Shiqi Chen","Tongyao Zhu","Ruochen Zhou","Jinghan Zhang","Siyang Gao","Juan Carlos Niebles","Mor Geva","Junxian He","Jiajun Wu","Manling Li"],"abstract":"Large Vision Language Models (VLMs) have long struggled with spatial reasoning tasks. Surprisingly, even simple spatial reasoning tasks, such as recognizing \"under\" or \"behind\" relationships between only two objects, pose significant challenges for current VLMs. In this work, we study the spatial reasoning challenge from the lens of mechanistic interpretability, diving into the model's internal states to examine the interactions between image and text tokens. By tracing attention distribution over the image through out intermediate layers, we observe that successful spatial reasoning correlates strongly with the model's ability to align its attention distribution with actual object locations, particularly differing between familiar and unfamiliar spatial relationships. Motivated by these findings, we propose ADAPTVIS based on inference-time confidence scores to sharpen the attention on highly relevant regions when confident, while smoothing and broadening the attention window to consider a wider context when confidence is lower. This training-free decoding method shows significant improvement (e.g., up to a 50 absolute point improvement) on spatial reasoning benchmarks such as WhatsUp and VSR with negligible cost. We make code and data publicly available for research purposes at https://github.com/shiqichen17/AdaptVis.","url_abs":"https://arxiv.org/abs/2503.01773v1","url_pdf":"https://arxiv.org/pdf/2503.01773v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"why-is-spatial-reasoning-hard-for-vlms-an","repo_url":"https://github.com/shiqichen17/adaptvis","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"spatial-reasoning","task_name":"Spatial Reasoning"}],"methods":[{"method_slug":"align","method_name":"ALIGN"},{"method_slug":"attention","method_name":"Attention"},{"method_slug":"softmax","method_name":"Softmax"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2503.01773","atlas_url":"https://app.syntology.ai/?focus=2503.01773","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2503.01773"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/shiqichen17/adaptvis","reach":null}],"summary":{"ran":5,"ran_draft_wrong":2,"unverified":1},"by_repo_kind":{"official":{"samples":7,"ran":6,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":8,"samples":[{"code_sha256_prefix":"3bb17da0b1fd87a0","entry":"LLaMAConfig","repo":"shiqichen17/adaptvis","repo_kind":"official","path":"model_zoo/llava/modeling_llava_scal.py","file_url":"https://github.com/shiqichen17/adaptvis/blob/HEAD/model_zoo/llava/modeling_llava_scal.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"3bb17da0b1fd87a0"}},{"code_sha256_prefix":"b6aba7799ae22451","entry":"LlavaCausalLMOutputWithPast","repo":"shiqichen17/adaptvis","repo_kind":"official","path":"model_zoo/llava/modeling_llava_scal.py","file_url":"https://github.com/shiqichen17/adaptvis/blob/HEAD/model_zoo/llava/modeling_llava_scal.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"b6aba7799ae22451"}},{"code_sha256_prefix":"bb877aca152f9d19","entry":"LlavaConfig","repo":"shiqichen17/adaptvis","repo_kind":"official","path":"model_zoo/llava/modeling_llava_scal.py","file_url":"https://github.com/shiqichen17/adaptvis/blob/HEAD/model_zoo/llava/modeling_llava_scal.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"bb877aca152f9d19"}},{"code_sha256_prefix":"9f4f556850f9c8ad","entry":"LlavaMultiModalProjector","repo":"shiqichen17/adaptvis","repo_kind":"official","path":"model_zoo/llava/modeling_llava_scal.py","file_url":"https://github.com/shiqichen17/adaptvis/blob/HEAD/model_zoo/llava/modeling_llava_scal.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"9f4f556850f9c8ad"}},{"code_sha256_prefix":"8312302b9693b681","entry":"LlavaPreTrainedModel","repo":"shiqichen17/adaptvis","repo_kind":"official","path":"model_zoo/llava/modeling_llava_scal.py","file_url":"https://github.com/shiqichen17/adaptvis/blob/HEAD/model_zoo/llava/modeling_llava_scal.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"8312302b9693b681"}},{"code_sha256_prefix":"373a7df152e5a1a5","entry":"apply_rotary_pos_emb","repo":"shiqichen17/AdaptVis","repo_kind":"official","path":"model_zoo/llama/modeling_llama.py","file_url":"https://github.com/shiqichen17/AdaptVis/blob/HEAD/model_zoo/llama/modeling_llama.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"373a7df152e5a1a5"}},{"code_sha256_prefix":"b99eea6376d1e212","entry":"rotate_half","repo":null,"repo_kind":null,"path":null,"file_url":null,"link_basis":"identical_code_first_harvested_elsewhere","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":null,"inline_ok":false,"mcp_get_code":{"code_sha256":"b99eea6376d1e212"}},{"code_sha256_prefix":"bfc3af8ebf837352","entry":"LlavaForConditionalGenerationScal","repo":"shiqichen17/adaptvis","repo_kind":"official","path":"model_zoo/llava/modeling_llava_scal.py","file_url":"https://github.com/shiqichen17/adaptvis/blob/HEAD/model_zoo/llava/modeling_llava_scal.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"bfc3af8ebf837352"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}