Browse State-of-the-Art › Robot Manipulation
Robot Manipulation
154 papers with code · 5 benchmarks · 12 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
5 leaderboard tables shown for this task, 5 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| CALVIN (19 rows) | DreamVLA | DreamVLA: A Vision-Language-Action Model Dreamed with... | code | Syntology ran 2 of 3 samples · 1 unverified | Compare |
| RLBench (18 rows) | EquAct | EquAct: An SE(3)-Equivariant Multi-Task Transformer for Open-Loop... | — | — | Compare |
| SimplerEnv-Google Robot (9 rows) | SoFar | SoFar: Language-Grounded Orientation Bridges Spatial Reasoning and... | code | — | Compare |
| MimicGen (7 rows) | SDP | SE(3)-Equivariant Diffusion Policy in Spherical Fourier Space | code | Syntology ran 7 of 14 samples · 7 unverified | Compare |
| SimplerEnv-Widow X (7 rows) | SoFar | SoFar: Language-Grounded Orientation Bridges Spatial Reasoning and... | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
12 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
3 subtasks in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 154 papers with code (430 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
13 Jun 2024 3 repositories listed Syntology ran 2 of 10 samples · 8 unverifiedLarge policies pretrained on a combination of Internet-scale vision-language data and diverse robot demonstrations have the potential to change how we teach robots new skills: rather than training new behaviors from…
-
20 Dec 2023 3 repositories listed Syntology ran 0 of 1 samples · 1 unverifiedIn this paper, we extend the scope of this effectiveness by showing that visual robot manipulation can significantly benefit from large-scale video generative pre-training.
-
18 Feb 2025 2 repositories listed Syntology ran 2 of 16 samples · 14 unverifiedWe present Magma, a foundation model that serves multimodal AI agentic tasks in both the digital and physical worlds.
-
18 Feb 2025 2 repositories listedSpatial intelligence is a critical component of embodied AI, promoting robots to understand and interact with their environments.
-
30 Jun 2023 2 repositories listed3D perceptual representations are well suited for robot manipulation as they easily encode occlusions and simplify spatial reasoning.
-
5 Jun 2023 2 repositories listed Syntology ran 1 of 6 samples · 5 unverifiedSpecifically, LIBERO highlights five key research topics in LLDM: 1) how to efficiently transfer declarative knowledge, procedural knowledge, or the mixture of both; 2) how to design effective policy architectures and…
-
6 Oct 2022 2 repositories listed Syntology ran 0 of 8 samples · 8 unverifiedWe show that a wide spectrum of robot manipulation tasks can be expressed with multimodal prompts, interleaving textual and visual tokens.
-
11 Sep 2022 2 repositories listed Syntology ran 0 of 11 samples · 11 unverifiedIn human environments, robots are expected to accomplish a variety of manipulation tasks given simple natural language instructions.
-
24 May 2022 2 repositories listed Syntology ran 3 of 3 samples · 0 unverified · 2 pointer-only (licence)Our intuition is that disagreement in learned reward model reflects uncertainty in tailored human feedback and could be useful for exploration.
-
13 Apr 2022 2 repositories listed Syntology ran 0 of 6 samples · 6 unverifiedWe have open-sourced our implementation to facilitate future research in learning to perform many complex manipulation skills in a row specified with natural language.
-
28 Jan 2022 2 repositories listedWe develop an end-to-end manipulation method based solely on detection and introduce Task-focused Few-shot Object Detection (TFOD) to learn new objects and settings.
-
3 Nov 2020 2 repositories listed3D scene representation for robot manipulation should capture three key object properties: permanency -- objects that become occluded over time continue to exist; amodal completeness -- objects have 3D occupancy, even…
-
16 Oct 2019 2 repositories listedIn order to exploit this idea, we introduce a framework whereby an object locomotion policy is initially obtained using a realistic physics simulator.
-
18 Sep 2018 2 repositories listedAutonomous robot manipulation involves estimating the translation and orientation of the object to be manipulated as a 6-degree-of-freedom (6D) pose.
-
31 Mar 2018 2 repositories listed Syntology ran 0 of 2 samples · 2 unverifiedEstimating the 6D pose of objects from images is an important problem in various applications such as robot manipulation and virtual reality.
-
6 Jul 2025 1 repository listed Syntology ran 2 of 3 samples · 1 unverified · 3 pointer-only (licence)Recent advances in vision-language-action (VLA) models have shown promise in integrating image generation with action prediction to improve generalization and reasoning in robot manipulation.
-
6 Jun 2025 1 repository listedWith the generated 3D object optical flow, we propose a flow-guided rendering mechanism, which renders the predicted final state and leverages GPT-4o to assess whether the predicted flow aligns with the task description.
-
20 May 2025 1 repository listed Syntology ran 1 of 13 samples · 12 unverifiedWorld models predict state transitions in response to actions and are increasingly developed across diverse modalities.
-
14 May 2025 1 repository listedBy contrast, in action generation, the dimensionality of the target is comparatively small, and only the image condition is high-dimensional.
-
13 May 2025 1 repository listed Syntology ran 1 of 3 samples · 2 unverifiedTo address these limitations, we propose FSD (From Seeing to Doing), a novel vision-language model that generates intermediate representations through spatial relationship reasoning, providing fine-grained guidance for…
-
9 May 2025 1 repository listed Syntology ran 1 of 5 samples · 4 unverifiedLearned from internet-scale videos, the generalist policy can be deployed to various robots through efficient latent action decoding.
-
6 May 2025 1 repository listed Syntology ran 0 of 12 samples · 12 unverifiedDual-system VLA (Vision-Language-Action) architectures have become a hot topic in embodied intelligence research, but there is a lack of sufficient open-source work for further performance analysis and optimization.
-
5 May 2025 1 repository listedIn this work, we present a vision-based approach for grasp verification to determine whether the robotic gripper has successfully grasped an object.
-
3 Apr 2025 1 repository listedRobot vision has greatly benefited from advancements in multimodal fusion techniques and vision-language models (VLMs).
-
31 Mar 2025 1 repository listed Syntology ran 0 of 5 samples · 5 unverifiedScalable and reproducible policy evaluation has been a long-standing challenge in robot learning.
-
25 Mar 2025 1 repository listed Syntology ran 2 of 3 samples · 1 unverifiedWhile recent vision-language-action models trained on diverse robot datasets exhibit promising generalization capabilities with limited in-domain data, their reliance on compact action heads to predict discretized or…
-
23 Mar 2025 1 repository listed Syntology ran 0 of 5 samples · 5 unverifiedCreating a physical digital twin of a real-world object has immense potential in robotics, content creation, and XR.
-
5 Mar 2025 1 repository listedThis can result in a significant increase in the time required for data collection, a lack of flexibility in scene setups, and a high level of complexity in the repetition of experiments.
-
2 Mar 2025 1 repository listed Syntology ran 2 of 5 samples · 3 unverifiedHowever, existing egocentric video representation learning methods mainly focus on aligning video representation with high-level narrations, overlooking the intricate dynamics between hands and objects.
-
20 Feb 2025 1 repository listedHumans possess a unified cognitive ability to perceive, comprehend, and interact with the physical world.
Syntology lines on 18 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections