{"url":"/dataset/the-colosseum","name":"The COLOSSEUM","full_name":"The COLOSSEUM: A Benchmark for Evaluating Generalization for Robotic Manipulation","description_markdown":"To realize effective large-scale, real-world robotic applications, we must evaluate how well our robot policies adapt to changes in environmental conditions. Unfortunately, a majority of studies evaluate robot performance in environments closely resembling or even identical to the training setup.\r\n\r\nWe present Colosseum, a novel simulation benchmark,with 20 diverse manipulation tasks, that enables systematical evaluation of models across 12 axes of environmental perturbations. These perturbations include changes in color, texture, and size of objects, table-tops, and backgrounds; we also vary lighting, distractors, and camera pose. Using Colosseum, we compare 4 state-of-the-art manipulation models to reveal that their success rate degrades between 30-50% across these perturbation factors.\r\n\r\nWhen multiple perturbations are applied in unison, the success rate degrades > 75%. We identify that changing the number of distractor objects, target object color, or lighting conditions are the perturbations that reduce model performance the most. To verify the ecological validity of our results, we show that our results in simulation are correlated (R2 = 0.614) to similar perturbations in real-world experiments. We open source code for others to use Colosseum, and also release code to 3D print the objects used to replicate the real-world perturbations. Ultimately, we hope that Colosseum will serve as a benchmark to identify modeling decisions that systematically improve generalization for manipulation.","description_withheld":null,"homepage":"https://robot-colosseum.github.io/","introduced_date":"2024-02-13","introduced_date_note":null,"introduced_by":{"paper":"/paper/the-colosseum-a-benchmark-for-evaluating","title":"THE COLOSSEUM: A Benchmark for Evaluating Generalization for Robotic Manipulation","first_author":"Wilbert Pumacay","url":null},"license":null,"modalities":[{"name":"Images","url":"/datasets/modality/images"},{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Robot Manipulation Generalization","url":"/task/robot-manipulation-generalization","datasets_with_task":"/datasets/task/robot-manipulation-generalization"}],"languages":[],"variants":["The COLOSSEUM"],"data_loaders":[],"num_papers_in_archive":13,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/robot-manipulation-generalization-on-the","task":"Robot Manipulation Generalization","dataset_variant":"The COLOSSEUM","rows":9,"metrics":["Average decrease average across all perturbations"],"first_row_in_archive_order":{"model":"RVT","paper":"/paper/0-1-deep-neural-networks-via-block-coordinate","metrics":{"Average decrease average across all perturbations":"-14.5"},"code_links":[]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/generative-image-as-action-models","title":"Generative Image as Action Models","date":"2024-07-10","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":1,"samples_ran":0,"samples_unverified":1,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/rvt-2-learning-precise-manipulation-from-few","title":"RVT-2: Learning Precise Manipulation from Few Demonstrations","date":"2024-06-12","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":15,"samples_ran":15,"samples_unverified":0,"pointer_only_for_licence":15,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/3d-diffuser-actor-policy-diffusion-with-3d","title":"3D Diffuser Actor: Policy Diffusion with 3D Scene Representations","date":"2024-02-18","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/learning-fine-grained-bimanual-manipulation","title":"Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware","date":"2023-04-23","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/perceiver-actor-a-multi-task-transformer-for","title":"Perceiver-Actor: A Multi-Task Transformer for Robotic Manipulation","date":"2022-09-12","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/0-1-deep-neural-networks-via-block-coordinate","title":"0/1 Deep Neural Networks via Block Coordinate Descent","date":"2022-06-19","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/r3m-a-universal-visual-representation-for","title":"R3M: A Universal Visual Representation for Robot Manipulation","date":"2022-03-23","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":6,"samples_ran":3,"samples_unverified":3,"pointer_only_for_licence":6,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/masked-visual-pre-training-for-motor-control","title":"Masked Visual Pre-training for Motor Control","date":"2022-03-11","rows_on_this_dataset":1,"code_links":1,"syntology":null}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":3,"samples_harvested":22,"samples_ran":18,"samples_unverified":4,"pointer_only_for_licence":21,"papers_with_no_sample_that_ran":1,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}