Datasets › The COLOSSEUM
The COLOSSEUM (The COLOSSEUM: A Benchmark for Evaluating Generalization for Robotic Manipulation)
To realize effective large-scale, real-world robotic applications, we must evaluate how well our robot policies adapt to changes in environmental conditions. Unfortunately, a majority of studies evaluate robot performance in environments closely resembling or even identical to the training setup.
We present Colosseum, a novel simulation benchmark,with 20 diverse manipulation tasks, that enables systematical evaluation of models across 12 axes of environmental perturbations. These perturbations include changes in color, texture, and size of objects, table-tops, and backgrounds; we also vary lighting, distractors, and camera pose. Using Colosseum, we compare 4 state-of-the-art manipulation models to reveal that their success rate degrades between 30-50% across these perturbation factors.
When multiple perturbations are applied in unison, the success rate degrades > 75%. We identify that changing the number of distractor objects, target object color, or lighting conditions are the perturbations that reduce model performance the most. To verify the ecological validity of our results, we show that our results in simulation are correlated (R2 = 0.614) to similar perturbations in real-world experiments. We open source code for others to use Colosseum, and also release code to 3D print the objects used to replicate the real-world perturbations. Ultimately, we hope that Colosseum will serve as a benchmark to identify modeling decisions that systematically improve generalization for manipulation.
Benchmarks archive 2025-07-28
All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.
| First row (archive order) | Paper | Code | ||||
|---|---|---|---|---|---|---|
| Robot Manipulation Generalization | The COLOSSEUM | RVT Average decrease average across all perturbations -14.5 | 0/1 Deep Neural Networks via Block Coordinate Descent | — | 9 | Compare |
Papers archive 2025-07-28
8 shown of 8 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 13. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.
| Date | Samples run Syntology | |||
|---|---|---|---|---|
| Generative Image as Action Models | 1 | 1 | 10 Jul 2024 | ran 0 of 1 samples (1 unverified) |
| RVT-2: Learning Precise Manipulation from Few Demonstrations | 1 | 1 | 12 Jun 2024 | ran 15 of 15 samples (0 unverified; 15 pointer-only for licence) |
| 3D Diffuser Actor: Policy Diffusion with 3D Scene Representations | 1 | 1 | 18 Feb 2024 | not harvested |
| Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware | 0 | 1 | 23 Apr 2023 | not harvested |
| Perceiver-Actor: A Multi-Task Transformer for Robotic Manipulation | 1 | 1 | 12 Sep 2022 | not harvested |
| 0/1 Deep Neural Networks via Block Coordinate Descent | 0 | 1 | 19 Jun 2022 | not harvested |
| R3M: A Universal Visual Representation for Robot Manipulation | 1 | 1 | 23 Mar 2022 | ran 3 of 6 samples (3 unverified; 6 pointer-only for licence) |
| Masked Visual Pre-training for Motor Control | 1 | 1 | 11 Mar 2022 | not harvested |
Dataset loaders archive 2025-07-28
No loader listed in the archive.
Tasks archive 2025-07-28
License archive 2025-07-28
No licence recorded in the archive. Absence here is not a statement about the dataset's terms.
Modalities archive 2025-07-28
Languages archive 2025-07-28
No language tagged.
Variants archive 2025-07-28
- The COLOSSEUM
1 variant name, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections