{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/unleashing-large-scale-video-generative-pre","title":"Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation","arxiv_id":"2312.13139","date":"2023-12-20","proceeding":null,"authors":["Hongtao Wu","Ya Jing","Chilam Cheang","Guangzeng Chen","Jiafeng Xu","Xinghang Li","Minghuan Liu","Hang Li","Tao Kong"],"abstract":"Generative pre-trained models have demonstrated remarkable effectiveness in language and vision domains by learning useful representations. In this paper, we extend the scope of this effectiveness by showing that visual robot manipulation can significantly benefit from large-scale video generative pre-training. We introduce GR-1, a straightforward GPT-style model designed for multi-task language-conditioned visual robot manipulation. GR-1 takes as inputs a language instruction, a sequence of observation images, and a sequence of robot states. It predicts robot actions as well as future images in an end-to-end manner. Thanks to a flexible design, GR-1 can be seamlessly finetuned on robot data after pre-trained on a large-scale video dataset. We perform extensive experiments on the challenging CALVIN benchmark and a real robot. On CALVIN benchmark, our method outperforms state-of-the-art baseline methods and improves the success rate from 88.9% to 94.9%. In the setting of zero-shot unseen scene generalization, GR-1 improves the success rate from 53.3% to 85.4%. In real robot experiments, GR-1 also outperforms baseline methods and shows strong potentials in generalization to unseen scenes and objects. We provide inaugural evidence that a unified GPT-style transformer, augmented with large-scale video generative pre-training, exhibits remarkable generalization to multi-task visual robot manipulation. Project page: https://GR1-Manipulation.github.io","url_abs":"https://arxiv.org/abs/2312.13139v2","url_pdf":"https://arxiv.org/pdf/2312.13139v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"unleashing-large-scale-video-generative-pre","repo_url":"https://github.com/bytedance/GR-MG","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"Apache-2.0"}},{"paper_slug":"unleashing-large-scale-video-generative-pre","repo_url":"https://github.com/bytedance/gr-1","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"unleashing-large-scale-video-generative-pre","repo_url":"https://github.com/GR1-Manipulation/GR-1","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"robot-manipulation","task_name":"Robot Manipulation"},{"task_slug":"zero-shot-generalization","task_name":"Zero-shot Generalization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/robot-manipulation-on-calvin","task":"Robot Manipulation","dataset":"CALVIN","model":"GR-1","rank_in_archive_order":15,"of":19,"metrics":{"avg. sequence length (D to D)":"3.06"},"uses_additional_data":false},{"leaderboard":"/sota/zero-shot-generalization-on-calvin","task":"Zero-shot Generalization","dataset":"CALVIN","model":"GR-1","rank_in_archive_order":5,"of":5,"metrics":{"Avg. sequence length":"3.06"},"uses_additional_data":true}],"syntology":{"syntology_url":"https://syntology.ai/paper/2312.13139","atlas_url":"https://app.syntology.ai/?focus=2312.13139","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2312.13139"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/GR1-Manipulation/GR-1","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/bytedance/GR-MG","reach":{"status":"ok","spdx":"Apache-2.0"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/bytedance/gr-1","reach":null}],"summary":{"unverified":1},"by_repo_kind":{"listed":{"samples":1,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"8e37b3c3fbc7c4a5","entry":"make_env","repo":"bytedance/gr-1","repo_kind":"listed","path":"evaluate_calvin.py","file_url":"https://github.com/bytedance/gr-1/blob/HEAD/evaluate_calvin.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"8e37b3c3fbc7c4a5"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}