{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/voxposer-composable-3d-value-maps-for-robotic","title":"VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models","arxiv_id":"2307.05973","date":"2023-07-12","proceeding":null,"authors":["Wenlong Huang","Chen Wang","Ruohan Zhang","Yunzhu Li","Jiajun Wu","Li Fei-Fei"],"abstract":"Large language models (LLMs) are shown to possess a wealth of actionable knowledge that can be extracted for robot manipulation in the form of reasoning and planning. Despite the progress, most still rely on pre-defined motion primitives to carry out the physical interactions with the environment, which remains a major bottleneck. In this work, we aim to synthesize robot trajectories, i.e., a dense sequence of 6-DoF end-effector waypoints, for a large variety of manipulation tasks given an open-set of instructions and an open-set of objects. We achieve this by first observing that LLMs excel at inferring affordances and constraints given a free-form language instruction. More importantly, by leveraging their code-writing capabilities, they can interact with a vision-language model (VLM) to compose 3D value maps to ground the knowledge into the observation space of the agent. The composed value maps are then used in a model-based planning framework to zero-shot synthesize closed-loop robot trajectories with robustness to dynamic perturbations. We further demonstrate how the proposed framework can benefit from online experiences by efficiently learning a dynamics model for scenes that involve contact-rich interactions. We present a large-scale study of the proposed method in both simulated and real-robot environments, showcasing the ability to perform a large variety of everyday manipulation tasks specified in free-form natural language. Videos and code at https://voxposer.github.io","url_abs":"https://arxiv.org/abs/2307.05973v2","url_pdf":"https://arxiv.org/pdf/2307.05973v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"voxposer-composable-3d-value-maps-for-robotic","repo_url":"https://github.com/huangwl18/voxposer","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"form","task_name":"Form"},{"task_slug":"language-modelling","task_name":"Language Modelling"},{"task_slug":"object-rearrangement","task_name":"Object Rearrangement"},{"task_slug":"robot-manipulation","task_name":"Robot Manipulation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/object-rearrangement-on-open6dor-v2","task":"Object Rearrangement","dataset":"Open6DOR V2","model":"VoxPoser","rank_in_archive_order":5,"of":5,"metrics":{"6-DoF":"-","pos-level0":"21.7","pos-level1":"35.6","rot-level0":"-","rot-level1":"-","rot-level2":"-"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=2307.05973","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2307.05973"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/huangwl18/voxposer","reach":null}],"summary":{"ran_draft_wrong":1},"by_repo_kind":{"official":{"samples":1,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"d807172220243c93","entry":"merge_dicts","repo":"huangwl18/voxposer","repo_kind":"official","path":"src/LMP.py","file_url":"https://github.com/huangwl18/voxposer/blob/HEAD/src/LMP.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"d807172220243c93"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}