{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-fine-grained-bimanual-manipulation","title":"Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware","arxiv_id":"2304.13705","date":"2023-04-23","proceeding":null,"authors":["Tony Z. Zhao","Vikash Kumar","Sergey Levine","Chelsea Finn"],"abstract":"Fine manipulation tasks, such as threading cable ties or slotting a battery, are notoriously difficult for robots because they require precision, careful coordination of contact forces, and closed-loop visual feedback. Performing these tasks typically requires high-end robots, accurate sensors, or careful calibration, which can be expensive and difficult to set up. Can learning enable low-cost and imprecise hardware to perform these fine manipulation tasks? We present a low-cost system that performs end-to-end imitation learning directly from real demonstrations, collected with a custom teleoperation interface. Imitation learning, however, presents its own challenges, particularly in high-precision domains: errors in the policy can compound over time, and human demonstrations can be non-stationary. To address these challenges, we develop a simple yet novel algorithm, Action Chunking with Transformers (ACT), which learns a generative model over action sequences. ACT allows the robot to learn 6 difficult tasks in the real world, such as opening a translucent condiment cup and slotting a battery with 80-90% success, with only 10 minutes worth of demonstrations. Project website: https://tonyzhaozh.github.io/aloha/","url_abs":"https://arxiv.org/abs/2304.13705v1","url_pdf":"https://arxiv.org/pdf/2304.13705v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"chunking","task_name":"Chunking"},{"task_slug":"imitation-learning","task_name":"Imitation Learning"},{"task_slug":"robot-manipulation","task_name":"Robot Manipulation"},{"task_slug":"robot-manipulation-generalization","task_name":"Robot Manipulation Generalization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/robot-manipulation-on-mimicgen","task":"Robot Manipulation","dataset":"MimicGen","model":"ACT (Evaluated in EquiDiff)","rank_in_archive_order":7,"of":7,"metrics":{"Succ. Rate (12 tasks, 100 demo/task)":"21.3","Succ. Rate (12 tasks, 1000 demo/task)":"63.3","Succ. Rate (12 tasks, 200 demo/task)":"38.2"},"uses_additional_data":false},{"leaderboard":"/sota/robot-manipulation-generalization-on-the","task":"Robot Manipulation Generalization","dataset":"The COLOSSEUM","model":"ACT","rank_in_archive_order":2,"of":9,"metrics":{"Average decrease average across all perturbations":"-61.8"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2304.13705","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}