{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-manipulation-by-predicting","title":"Learning Manipulation by Predicting Interaction","arxiv_id":"2406.00439","date":"2024-06-01","proceeding":null,"authors":["Jia Zeng","Qingwen Bu","Bangjun Wang","Wenke Xia","Li Chen","Hao Dong","Haoming Song","Dong Wang","Di Hu","Ping Luo","Heming Cui","Bin Zhao","Xuelong Li","Yu Qiao","Hongyang Li"],"abstract":"Representation learning approaches for robotic manipulation have boomed in recent years. Due to the scarcity of in-domain robot data, prevailing methodologies tend to leverage large-scale human video datasets to extract generalizable features for visuomotor policy learning. Despite the progress achieved, prior endeavors disregard the interactive dynamics that capture behavior patterns and physical interaction during the manipulation process, resulting in an inadequate understanding of the relationship between objects and the environment. To this end, we propose a general pre-training pipeline that learns Manipulation by Predicting the Interaction (MPI) and enhances the visual representation.Given a pair of keyframes representing the initial and final states, along with language instructions, our algorithm predicts the transition frame and detects the interaction object, respectively. These two learning objectives achieve superior comprehension towards \"how-to-interact\" and \"where-to-interact\". We conduct a comprehensive evaluation of several challenging robotic tasks.The experimental results demonstrate that MPI exhibits remarkable improvement by 10% to 64% compared with previous state-of-the-art in real-world robot platforms as well as simulation environments. Code and checkpoints are publicly shared at https://github.com/OpenDriveLab/MPI.","url_abs":"https://arxiv.org/abs/2406.00439v1","url_pdf":"https://arxiv.org/pdf/2406.00439v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"learning-manipulation-by-predicting","repo_url":"https://github.com/opendrivelab/mpi","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"representation-learning","task_name":"Representation Learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2406.00439","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2406.00439"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/opendrivelab/mpi","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"ran":3},"by_repo_kind":{"official":{"samples":3,"ran":3,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"38df077a6817f796","entry":"load_json_file","repo":"opendrivelab/mpi","repo_kind":"official","path":"mpi/datasets/ego4d_hoi_dataset.py","file_url":"https://github.com/opendrivelab/mpi/blob/HEAD/mpi/datasets/ego4d_hoi_dataset.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"38df077a6817f796"}},{"code_sha256_prefix":"259b0af842be3b00","entry":"load_pickle","repo":"opendrivelab/mpi","repo_kind":"official","path":"mpi/datasets/ego4d_hoi_dataset.py","file_url":"https://github.com/opendrivelab/mpi/blob/HEAD/mpi/datasets/ego4d_hoi_dataset.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"259b0af842be3b00"}},{"code_sha256_prefix":"875c25be90cdd3b5","entry":"scaled_center_crop","repo":"opendrivelab/mpi","repo_kind":"official","path":"mpi/datasets/ego4d_hoi_dataset.py","file_url":"https://github.com/opendrivelab/mpi/blob/HEAD/mpi/datasets/ego4d_hoi_dataset.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"875c25be90cdd3b5"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}