Datasets › CALVIN

CALVIN (Composing Actions from Language and Vision)

Introduced by Oier Mees et al. in CALVIN: A Benchmark for Language-Conditioned Policy Learning for Long-Horizon Robot Manipulation Tasks6 Dec 2021 archive 2025-07-28

CALVIN (Composing Actions from Language and Vision), is an open-source simulated benchmark to learn long-horizon language-conditioned robot manipulation tasks.

Benchmarks archive 2025-07-28

All 2 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Robot Manipulation CALVIN DreamVLA avg. sequence length (D to D) 4.44 DreamVLA: A Vision-Language-Action Model Dreamed with... Zhangwenyao1/DreamVLA 19 Compare
Zero-shot Generalization CALVIN GR-MG Avg. sequence length 4.04 GR-MG: Leveraging Partially Annotated Data via... bytedance/GR-MG 5 Compare

Papers archive 2025-07-28

18 shown of 18 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 65. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge 1 1 6 Jul 2025 ran 2 of 3 samples (1 unverified; 3 pointer-only for licence)
UniVLA: Learning to Act Anywhere with Task-centric Latent Actions 1 1 9 May 2025 ran 1 of 5 samples (4 unverified)
OpenHelix: A Short Survey, Empirical Analysis, and Open-Source Dual-System VLA Model for Robotic Manipulation 1 1 6 May 2025 ran 0 of 12 samples (12 unverified)
UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent 0 1 31 Jan 2025 not harvested
Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations 1 1 19 Dec 2024 ran 4 of 5 samples (1 unverified)
Towards Generalist Robot Policies: What Matters in Building Vision-Language-Action Models 1 1 18 Dec 2024 ran 1 of 8 samples (7 unverified)
Efficient Diffusion Transformer Policies with Mixture of Expert Denoisers for Multitask Learning 1 2 17 Dec 2024 ran 5 of 11 samples (6 unverified)
VidMan: Exploiting Implicit Dynamics from Video Diffusion Model for Effective Robot Manipulation 0 1 14 Nov 2024 not harvested
Towards Synergistic, Generalized, and Efficient Dual-System for Robotic Manipulation 0 1 10 Oct 2024 not harvested
GR-MG: Leveraging Partially Annotated Data via Multi-Modal Goal-Conditioned Policy 1 2 26 Aug 2024 ran 8 of 10 samples (2 unverified)
RoboUniView: Visual-Language Model with Unified View Representation for Robotic Manipulation 1 2 27 Jun 2024 ran 1 of 1 samples (0 unverified; 1 pointer-only for licence)
OpenVLA: An Open-Source Vision-Language-Action Model 3 1 13 Jun 2024 ran 2 of 10 samples (8 unverified)
From LLMs to Actions: Latent Codes as Bridges in Hierarchical Robot Control 0 1 8 May 2024 not harvested
3D Diffuser Actor: Policy Diffusion with 3D Scene Representations 1 3 18 Feb 2024 not harvested
Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation 3 2 20 Dec 2023 ran 0 of 1 samples (1 unverified)
Vision-Language Foundation Models as Effective Robot Imitators 0 1 2 Nov 2023 not harvested
Learning Universal Policies via Text-Guided Video Generation 0 1 31 Jan 2023 not harvested
RT-1: Robotics Transformer for Real-World Control at Scale 1 1 13 Dec 2022 ran 0 of 4 samples (4 unverified)

Dataset loaders archive 2025-07-28

1 loader as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

MIT License

Modalities archive 2025-07-28

No modality tagged.

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • CALVIN

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections