{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/thinking-in-space-how-multimodal-large","title":"Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces","arxiv_id":"2412.14171","date":"2024-12-18","proceeding":"CVPR 2025 1","authors":["Jihan Yang","Shusheng Yang","Anjali W. Gupta","Rilyn Han","Li Fei-Fei","Saining Xie"],"abstract":"Humans possess the visual-spatial intelligence to remember spaces from sequential visual observations. However, can Multimodal Large Language Models (MLLMs) trained on million-scale video datasets also ``think in space'' from videos? We present a novel video-based visual-spatial intelligence benchmark (VSI-Bench) of over 5,000 question-answer pairs, and find that MLLMs exhibit competitive - though subhuman - visual-spatial intelligence. We probe models to express how they think in space both linguistically and visually and find that while spatial reasoning capabilities remain the primary bottleneck for MLLMs to reach higher benchmark performance, local world models and spatial awareness do emerge within these models. Notably, prevailing linguistic reasoning techniques (e.g., chain-of-thought, self-consistency, tree-of-thoughts) fail to improve performance, whereas explicitly generating cognitive maps during question-answering enhances MLLMs' spatial distance ability.","url_abs":"https://arxiv.org/abs/2412.14171v1","url_pdf":"https://arxiv.org/pdf/2412.14171v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"thinking-in-space-how-multimodal-large","repo_url":"https://github.com/vision-x-nyu/thinking-in-space","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":"question-answering","task_name":"Question Answering"},{"task_slug":"spatial-reasoning","task_name":"Spatial Reasoning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2412.14171","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2412.14171"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/vision-x-nyu/thinking-in-space","reach":{"status":"ok","spdx":"Apache-2.0"}}],"summary":{"ran":2,"ran_draft_wrong":2,"unverified":1},"by_repo_kind":{"official":{"samples":5,"ran":4,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"00f5e7d9ee6636ba","entry":"get_sample_size","repo":"vision-x-nyu/thinking-in-space","repo_kind":"official","path":"lmms_eval/evaluator_utils.py","file_url":"https://github.com/vision-x-nyu/thinking-in-space/blob/HEAD/lmms_eval/evaluator_utils.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"00f5e7d9ee6636ba"}},{"code_sha256_prefix":"20a7cc804eb22661","entry":"hash_args","repo":"vision-x-nyu/thinking-in-space","repo_kind":"official","path":"lmms_eval/api/model.py","file_url":"https://github.com/vision-x-nyu/thinking-in-space/blob/HEAD/lmms_eval/api/model.py","link_basis":"harvester_set","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":true,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"20a7cc804eb22661"}},{"code_sha256_prefix":"74386ef186ff5c64","entry":"remove_none_pattern","repo":"vision-x-nyu/thinking-in-space","repo_kind":"official","path":"lmms_eval/logging_utils.py","file_url":"https://github.com/vision-x-nyu/thinking-in-space/blob/HEAD/lmms_eval/logging_utils.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"74386ef186ff5c64"}},{"code_sha256_prefix":"c44830a7722ff5f7","entry":"request_caching_arg_to_dict","repo":"vision-x-nyu/thinking-in-space","repo_kind":"official","path":"lmms_eval/evaluator.py","file_url":"https://github.com/vision-x-nyu/thinking-in-space/blob/HEAD/lmms_eval/evaluator.py","link_basis":"plan_row","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":true,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"c44830a7722ff5f7"}},{"code_sha256_prefix":"72b68ba343decfd1","entry":"calculate_sha256","repo":"vision-x-nyu/thinking-in-space","repo_kind":"official","path":"lmms_eval/models/gpt4v.py","file_url":"https://github.com/vision-x-nyu/thinking-in-space/blob/HEAD/lmms_eval/models/gpt4v.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"72b68ba343decfd1"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}