{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/large-spatial-model-end-to-end-unposed-images","title":"Large Spatial Model: End-to-end Unposed Images to Semantic 3D","arxiv_id":"2410.18956","date":"2024-10-24","proceeding":null,"authors":["Zhiwen Fan","Jian Zhang","Wenyan Cong","Peihao Wang","Renjie Li","Kairun Wen","Shijie Zhou","Achuta Kadambi","Zhangyang Wang","Danfei Xu","Boris Ivanovic","Marco Pavone","Yue Wang"],"abstract":"Reconstructing and understanding 3D structures from a limited number of images is a well-established problem in computer vision. Traditional methods usually break this task into multiple subtasks, each requiring complex transformations between different data representations. For instance, dense reconstruction through Structure-from-Motion (SfM) involves converting images into key points, optimizing camera parameters, and estimating structures. Afterward, accurate sparse reconstructions are required for further dense modeling, which is subsequently fed into task-specific neural networks. This multi-step process results in considerable processing time and increased engineering complexity. In this work, we present the Large Spatial Model (LSM), which processes unposed RGB images directly into semantic radiance fields. LSM simultaneously estimates geometry, appearance, and semantics in a single feed-forward operation, and it can generate versatile label maps by interacting with language at novel viewpoints. Leveraging a Transformer-based architecture, LSM integrates global geometry through pixel-aligned point maps. To enhance spatial attribute regression, we incorporate local context aggregation with multi-scale fusion, improving the accuracy of fine local details. To tackle the scarcity of labeled 3D semantic data and enable natural language-driven scene manipulation, we incorporate a pre-trained 2D language-based segmentation model into a 3D-consistent semantic feature field. An efficient decoder then parameterizes a set of semantic anisotropic Gaussians, facilitating supervised end-to-end learning. Extensive experiments across various tasks show that LSM unifies multiple 3D vision tasks directly from unposed images, achieving real-time semantic 3D reconstruction for the first time.","url_abs":"https://arxiv.org/abs/2410.18956v2","url_pdf":"https://arxiv.org/pdf/2410.18956v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"large-spatial-model-end-to-end-unposed-images","repo_url":"https://github.com/NVlabs/LSM","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"NOASSERTION"}}],"tasks":[{"task_slug":"3d-reconstruction","task_name":"3D Reconstruction"},{"task_slug":"attribute","task_name":"Attribute"}],"methods":[{"method_slug":"set","method_name":"SET"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2410.18956","atlas_url":"https://app.syntology.ai/?focus=2410.18956","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2410.18956"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/NVlabs/LSM","reach":{"status":"ok","spdx":"NOASSERTION"}}],"summary":{"unverified":4},"by_repo_kind":{"official":{"samples":4,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":4,"samples":[{"code_sha256_prefix":"29da0adf4a5d0c31","entry":"calculate_depth_metrics","repo":"NVlabs/LSM","repo_kind":"official","path":"large_spatial_model/loss.py","file_url":"https://github.com/NVlabs/LSM/blob/HEAD/large_spatial_model/loss.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"29da0adf4a5d0c31"}},{"code_sha256_prefix":"a66fba2aa76f63ec","entry":"forward_layers","repo":"NVlabs/LSM","repo_kind":"official","path":"large_spatial_model/lseg.py","file_url":"https://github.com/NVlabs/LSM/blob/HEAD/large_spatial_model/lseg.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"a66fba2aa76f63ec"}},{"code_sha256_prefix":"b78a32ba60c47f45","entry":"map_func","repo":"NVlabs/LSM","repo_kind":"official","path":"large_spatial_model/datasets/testdata.py","file_url":"https://github.com/NVlabs/LSM/blob/HEAD/large_spatial_model/datasets/testdata.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"b78a32ba60c47f45"}},{"code_sha256_prefix":"1c27cad2381d6172","entry":"merge_and_split_predictions","repo":"NVlabs/LSM","repo_kind":"official","path":"large_spatial_model/loss.py","file_url":"https://github.com/NVlabs/LSM/blob/HEAD/large_spatial_model/loss.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"1c27cad2381d6172"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}