{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/language-conditioned-spatial-relation","title":"Language Conditioned Spatial Relation Reasoning for 3D Object Grounding","arxiv_id":"2211.09646","date":"2022-11-17","proceeding":null,"authors":["ShiZhe Chen","Pierre-Louis Guhur","Makarand Tapaswi","Cordelia Schmid","Ivan Laptev"],"abstract":"Localizing objects in 3D scenes based on natural language requires understanding and reasoning about spatial relations. In particular, it is often crucial to distinguish similar objects referred by the text, such as \"the left most chair\" and \"a chair next to the window\". In this work we propose a language-conditioned transformer model for grounding 3D objects and their spatial relations. To this end, we design a spatial self-attention layer that accounts for relative distances and orientations between objects in input 3D point clouds. Training such a layer with visual and language inputs enables to disambiguate spatial relations and to localize objects referred by the text. To facilitate the cross-modal learning of relations, we further propose a teacher-student approach where the teacher model is first trained using ground-truth object labels, and then helps to train a student model using point cloud inputs. We perform ablation studies showing advantages of our approach. We also demonstrate our model to significantly outperform the state of the art on the challenging Nr3D, Sr3D and ScanRefer 3D object grounding datasets.","url_abs":"https://arxiv.org/abs/2211.09646v1","url_pdf":"https://arxiv.org/pdf/2211.09646v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"language-conditioned-spatial-relation","repo_url":"https://github.com/cshizhe/vil3dref","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"object","task_name":"Object"},{"task_slug":null,"task_name":"Relation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2211.09646","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2211.09646"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/cshizhe/vil3dref","reach":null}],"summary":{"ran":2,"unverified":2},"by_repo_kind":{"official":{"samples":4,"ran":2,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":4,"samples":[{"code_sha256_prefix":"358ccd9d48a67d8c","entry":"MultiHeadAttentionSpatial","repo":"cshizhe/vil3dref","repo_kind":"official","path":"og3d_src/model/cmt_module.py","file_url":"https://github.com/cshizhe/vil3dref/blob/HEAD/og3d_src/model/cmt_module.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"358ccd9d48a67d8c"}},{"code_sha256_prefix":"826240c2b4271a47","entry":"TransformerDecoderLayer","repo":"cshizhe/vil3dref","repo_kind":"official","path":"og3d_src/model/cmt_module.py","file_url":"https://github.com/cshizhe/vil3dref/blob/HEAD/og3d_src/model/cmt_module.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"826240c2b4271a47"}},{"code_sha256_prefix":"22728ff50a36e809","entry":"CMT","repo":"cshizhe/vil3dref","repo_kind":"official","path":"og3d_src/model/cmt_module.py","file_url":"https://github.com/cshizhe/vil3dref/blob/HEAD/og3d_src/model/cmt_module.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"22728ff50a36e809"}},{"code_sha256_prefix":"784b909829e3a003","entry":"TransformerSpatialDecoderLayer","repo":"cshizhe/vil3dref","repo_kind":"official","path":"og3d_src/model/cmt_module.py","file_url":"https://github.com/cshizhe/vil3dref/blob/HEAD/og3d_src/model/cmt_module.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"784b909829e3a003"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}