{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/sofar-language-grounded-orientation-bridges","title":"SoFar: Language-Grounded Orientation Bridges Spatial Reasoning and Object Manipulation","arxiv_id":"2502.13143","date":"2025-02-18","proceeding":null,"authors":["Zekun Qi","Wenyao Zhang","Yufei Ding","Runpei Dong","Xinqiang Yu","Jingwen Li","Lingyun Xu","Baoyu Li","Xialin He","Guofan Fan","Jiazhao Zhang","JiaWei He","Jiayuan Gu","Xin Jin","Kaisheng Ma","Zhizheng Zhang","He Wang","Li Yi"],"abstract":"Spatial intelligence is a critical component of embodied AI, promoting robots to understand and interact with their environments. While recent advances have enhanced the ability of VLMs to perceive object locations and positional relationships, they still lack the capability to precisely understand object orientations-a key requirement for tasks involving fine-grained manipulations. Addressing this limitation not only requires geometric reasoning but also an expressive and intuitive way to represent orientation. In this context, we propose that natural language offers a more flexible representation space than canonical frames, making it particularly suitable for instruction-following robotic systems. In this paper, we introduce the concept of semantic orientation, which defines object orientations using natural language in a reference-frame-free manner (e.g., the ''plug-in'' direction of a USB or the ''handle'' direction of a knife). To support this, we construct OrienText300K, a large-scale dataset of 3D models annotated with semantic orientations that link geometric understanding to functional semantics. By integrating semantic orientation into a VLM system, we enable robots to generate manipulation actions with both positional and orientational constraints. Extensive experiments in simulation and real world demonstrate that our approach significantly enhances robotic manipulation capabilities, e.g., 48.7% accuracy on Open6DOR and 74.9% accuracy on SIMPLER.","url_abs":"https://arxiv.org/abs/2502.13143v1","url_pdf":"https://arxiv.org/pdf/2502.13143v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"sofar-language-grounded-orientation-bridges","repo_url":"https://github.com/qizekun/SoFar","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"sofar-language-grounded-orientation-bridges","repo_url":"https://github.com/zhangwenyao1/open6dor_v2_execution","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"object-rearrangement","task_name":"Object Rearrangement"},{"task_slug":"robot-manipulation","task_name":"Robot Manipulation"},{"task_slug":"robot-navigation","task_name":"Robot Navigation"},{"task_slug":"spatial-reasoning","task_name":"Spatial Reasoning"},{"task_slug":"visual-question-answering-1","task_name":"Visual Question Answering"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/object-rearrangement-on-open6dor-v2","task":"Object Rearrangement","dataset":"Open6DOR V2","model":"SoFar","rank_in_archive_order":1,"of":5,"metrics":{"6-DoF":"48.7","pos-level0":"96.0","pos-level1":"81.5","rot-level0":"68.6","rot-level1":"42.2","rot-level2":"70.1"},"uses_additional_data":false},{"leaderboard":"/sota/robot-manipulation-on-simpler-env","task":"Robot Manipulation","dataset":"SimplerEnv-Google Robot","model":"SoFar","rank_in_archive_order":1,"of":9,"metrics":{"Variant Aggregation":"0.676","Variant Aggregation-Move Near":"0.740","Variant Aggregation-Open/Close Drawer":"0.297","Variant Aggregation-Pick Coke Can":"0.907","Visual Matching":"0.749","Visual Matching-Move Near":"0.917","Visual Matching-Open/Close Drawer":"0.403","Visual Matching-Pick Coke Can":"0.923"},"uses_additional_data":false},{"leaderboard":"/sota/robot-manipulation-on-simplerenv-widow-x","task":"Robot Manipulation","dataset":"SimplerEnv-Widow X","model":"SoFar","rank_in_archive_order":1,"of":7,"metrics":{"Average":"0.583","Put Carrot on Plate":"0.667","Put Eggplant in Yellow Basket":"0.375","Put Spoon on Towel":"0.583","Stack Green Block on Yellow Block":"0.708"},"uses_additional_data":false},{"leaderboard":"/sota/spatial-reasoning-on-6-dof-spatialbench","task":"Spatial Reasoning","dataset":"6-DoF SpatialBench","model":"SoFar","rank_in_archive_order":1,"of":7,"metrics":{"Orientation-abs":"31.3","Orientation-rel":"54.6","Position-abs":"33.8","Position-rel":"59.6","Total":"43.9"},"uses_additional_data":false},{"leaderboard":"/sota/spatial-reasoning-on-embspatial-bench","task":"Spatial Reasoning","dataset":"EmbSpatial-Bench","model":"SoFar","rank_in_archive_order":1,"of":5,"metrics":{"Generation":"70.88"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2502.13143","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}