{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/clevrskills-compositional-language-and-visual","title":"ClevrSkills: Compositional Language and Visual Reasoning in Robotics","arxiv_id":"2411.09052","date":"2024-11-13","proceeding":null,"authors":["Sanjay Haresh","Daniel Dijkman","Apratim Bhattacharyya","Roland Memisevic"],"abstract":"Robotics tasks are highly compositional by nature. For example, to perform a high-level task like cleaning the table a robot must employ low-level capabilities of moving the effectors to the objects on the table, pick them up and then move them off the table one-by-one, while re-evaluating the consequently dynamic scenario in the process. Given that large vision language models (VLMs) have shown progress on many tasks that require high level, human-like reasoning, we ask the question: if the models are taught the requisite low-level capabilities, can they compose them in novel ways to achieve interesting high-level tasks like cleaning the table without having to be explicitly taught so? To this end, we present ClevrSkills - a benchmark suite for compositional reasoning in robotics. ClevrSkills is an environment suite developed on top of the ManiSkill2 simulator and an accompanying dataset. The dataset contains trajectories generated on a range of robotics tasks with language and visual annotations as well as multi-modal prompts as task specification. The suite includes a curriculum of tasks with three levels of compositional understanding, starting with simple tasks requiring basic motor skills. We benchmark multiple different VLM baselines on ClevrSkills and show that even after being pre-trained on large numbers of tasks, these models fail on compositional reasoning in robotics tasks.","url_abs":"https://arxiv.org/abs/2411.09052v1","url_pdf":"https://arxiv.org/pdf/2411.09052v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"clevrskills-compositional-language-and-visual","repo_url":"https://github.com/Qualcomm-AI-research/ClevrSkills","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"visual-reasoning","task_name":"Visual Reasoning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2411.09052","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2411.09052"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/Qualcomm-AI-research/ClevrSkills","reach":null}],"summary":{"ran_violates":1,"ran_draft_wrong":1,"unverified":1},"by_repo_kind":{"official":{"samples":3,"ran":2,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":3,"samples":[{"code_sha256_prefix":"c1bb265a27408726","entry":"discretize_actions","repo":"Qualcomm-AI-research/ClevrSkills","repo_kind":"official","path":"clevr_skills/clevr_skills_oracle.py","file_url":"https://github.com/Qualcomm-AI-research/ClevrSkills/blob/HEAD/clevr_skills/clevr_skills_oracle.py","link_basis":"first_harvest_node","language":"python","status":"ran_violates","verification_level":1,"contract_check":"VIOLATES","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":"BSD-3-Clause-Clear","inline_ok":false,"mcp_get_code":{"code_sha256":"c1bb265a27408726"}},{"code_sha256_prefix":"af392e91c9565f8a","entry":"get_render_config","repo":"Qualcomm-AI-research/ClevrSkills","repo_kind":"official","path":"clevr_skills/clevr_skills_oracle.py","file_url":"https://github.com/Qualcomm-AI-research/ClevrSkills/blob/HEAD/clevr_skills/clevr_skills_oracle.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"BSD-3-Clause-Clear","inline_ok":false,"mcp_get_code":{"code_sha256":"af392e91c9565f8a"}},{"code_sha256_prefix":"19f36f7584839e81","entry":"get_action_trace","repo":"Qualcomm-AI-research/ClevrSkills","repo_kind":"official","path":"clevr_skills/clevr_skills_oracle.py","file_url":"https://github.com/Qualcomm-AI-research/ClevrSkills/blob/HEAD/clevr_skills/clevr_skills_oracle.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"BSD-3-Clause-Clear","inline_ok":false,"mcp_get_code":{"code_sha256":"19f36f7584839e81"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}