Papers › Perceiver-Actor: A Multi-Task Transformer for Robotic Manipulation

Perceiver-Actor: A Multi-Task Transformer for Robotic Manipulation

12 Sep 2022arXiv:2209.05451archive 2025-07-28

Mohit Shridhar, Lucas Manuelli, Dieter Fox

Transformers have revolutionized vision and natural language processing with their ability to scale with large datasets. But in robotic manipulation, data is both limited and expensive. Can manipulation still benefit from Transformers with the right problem formulation? We investigate this question with PerAct, a language-conditioned behavior-cloning agent for multi-task 6-DoF manipulation. PerAct encodes language goals and RGB-D voxel observations with a Perceiver Transformer, and outputs discretized actions by ``detecting the next best voxel action''. Unlike frameworks that operate on 2D images, the voxelized 3D observation and action space provides a strong structural prior for efficiently learning 6-DoF actions. With this formulation, we train a single multi-task Transformer for 18 RLBench tasks (with 249 variations) and 7 real-world tasks (with 18 variations) from just a few demonstrations per task. Our results show that PerAct significantly outperforms unstructured image-to-action agents and 3D ConvNet baselines for a wide range of tabletop tasks.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

peract/peract officialmentioned on GitHubpytorchApache-2.0 report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Robot ManipulationRobot Manipulation Generalization

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Robot Manipulation RLBench PerAct (Evaluated in RVT) Inference Speed (fps) 4.9 #11 of 18 Archive leaderboard report
Robot Manipulation RLBench PerAct (Evaluated in RVT) Input Image Size 128 #11 of 18 Archive leaderboard report
Robot Manipulation RLBench PerAct (Evaluated in RVT) Succ. Rate (18 tasks, 100 demo/task) 49.4 #11 of 18 Archive leaderboard report
Robot Manipulation RLBench PerAct (Evaluated in RVT) Training Time (V100 x 8 x day) 16 #11 of 18 Archive leaderboard report
Robot Manipulation RLBench PerAct Input Image Size 128 #14 of 18 Archive leaderboard report
Robot Manipulation RLBench PerAct Succ. Rate (18 tasks, 10 demo/task) 30 #14 of 18 Archive leaderboard report
Robot Manipulation RLBench PerAct Succ. Rate (18 tasks, 100 demo/task) 42.7 #14 of 18 Archive leaderboard report
Robot Manipulation RLBench PerAct Training Time (V100 x 8 x day) 16 #14 of 18 Archive leaderboard report
Robot Manipulation RLBench Image-BC VIT Input Image Size 128 #16 of 18 Archive leaderboard report
Robot Manipulation RLBench Image-BC VIT Succ. Rate (18 tasks, 100 demo/task) 1.3 #16 of 18 Archive leaderboard report
Robot Manipulation RLBench Image-BC CNN Input Image Size 128 #17 of 18 Archive leaderboard report
Robot Manipulation RLBench Image-BC CNN Succ. Rate (18 tasks, 100 demo/task) 1.3 #17 of 18 Archive leaderboard report
Robot Manipulation Generalization The COLOSSEUM PerAct Average decrease average across all perturbations -17.3 #6 of 9 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Absolute Position EncodingsAdamAttentionBPEDense ConnectionsDropoutLabel SmoothingLayer NormalizationLinear LayerMulti-Head AttentionPosition-Wise Feed-Forward LayerResidual ConnectionSoftmaxTransformer

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections