{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/coarse-to-fine-q-attention-efficient-learning","title":"Coarse-to-Fine Q-attention: Efficient Learning for Visual Robotic Manipulation via Discretisation","arxiv_id":"2106.12534","date":"2021-06-23","proceeding":"CVPR 2022 1","authors":["Stephen James","Kentaro Wada","Tristan Laidlow","Andrew J. Davison"],"abstract":"We present a coarse-to-fine discretisation method that enables the use of discrete reinforcement learning approaches in place of unstable and data-inefficient actor-critic methods in continuous robotics domains. This approach builds on the recently released ARM algorithm, which replaces the continuous next-best pose agent with a discrete one, with coarse-to-fine Q-attention. Given a voxelised scene, coarse-to-fine Q-attention learns what part of the scene to 'zoom' into. When this 'zooming' behaviour is applied iteratively, it results in a near-lossless discretisation of the translation space, and allows the use of a discrete action, deep Q-learning method. We show that our new coarse-to-fine algorithm achieves state-of-the-art performance on several difficult sparsely rewarded RLBench vision-based robotics tasks, and can train real-world policies, tabula rasa, in a matter of minutes, with as little as 3 demonstrations.","url_abs":"https://arxiv.org/abs/2106.12534v2","url_pdf":"https://arxiv.org/pdf/2106.12534v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"coarse-to-fine-q-attention-efficient-learning","repo_url":"https://github.com/stepjam/ARM","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"NOASSERTION"}}],"tasks":[{"task_slug":"continuous-control","task_name":"Continuous Control"},{"task_slug":"q-learning","task_name":"Q-Learning"},{"task_slug":"robot-manipulation","task_name":"Robot Manipulation"},{"task_slug":"translation","task_name":"Translation"}],"methods":[{"method_slug":"q-learning","method_name":"Q-Learning"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/robot-manipulation-on-rlbench","task":"Robot Manipulation","dataset":"RLBench","model":"C2FARM-BC (Evaluated in PerAct)","rank_in_archive_order":15,"of":18,"metrics":{"Input Image Size":"128","Succ. Rate (18 tasks, 100 demo/task)":"20.1"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/2106.12534","atlas_url":"https://app.syntology.ai/?focus=2106.12534","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}