Datasets › BAIR Robot Pushing

BAIR Robot Pushing

archive 2025-07-28

Dataset of 64x64 images of a robot pushing objects on a table top. From Berkeley AI Research (BAIR).

Source: Self-Supervised Visual Planning with Temporal Skip Connections (https://arxiv.org/abs/1710.05268)

Video prediction : Conditioned on 2 frames, predict 14 frames.

Benchmarks archive 2025-07-28

All 2 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Video Generation BAIR Robot Pushing MAGVIT FVD score 62 MAGVIT: Masked Generative Video Transformer google-research/magvit 31 Compare
Video Prediction BAIR Robot Pushing MAGVIT (-L-FP) FVD 62±0.1 MAGVIT: Masked Generative Video Transformer google-research/magvit 6 Compare

Papers archive 2025-07-28

24 shown of 24 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 27. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
MAGVIT: Masked Generative Video Transformer 1 3 10 Dec 2022 ran 1 of 10 samples (9 unverified)
Tell Me What Happened: Unifying Text-guided Video Completion via Multimodal Masked Video Generation 1 1 23 Nov 2022 not harvested
Phenaki: Variable Length Video Generation From Open Domain Textual Description 2 1 5 Oct 2022 ran 3 of 3 samples (0 unverified; 2 pointer-only for licence)
Diffusion Models for Video Prediction and Infilling 1 1 15 Jun 2022 ran 9 of 19 samples (10 unverified)
MCVD: Masked Conditional Video Diffusion for Prediction, Generation, and Interpolation 2 3 19 May 2022 ran 7 of 16 samples (9 unverified)
NÜWA: Visual Synthesis Pre-training for Neural visUal World creAtion 1 1 24 Nov 2021 ran 3 of 3 samples (0 unverified; 2 pointer-only for licence)
SLAMP: Stochastic Latent Appearance and Motion Prediction 1 1 5 Aug 2021 ran 2 of 2 samples (0 unverified)
CCVS: Context-aware Controllable Video Synthesis 1 1 16 Jul 2021 ran 3 of 8 samples (5 unverified)
Diverse Video Generation using a Gaussian Process Trigger 1 1 9 Jul 2021 not harvested
FitVid: Overfitting in Pixel-Level Video Prediction 1 1 24 Jun 2021 ran 0 of 11 samples (11 unverified)
VideoGPT: Video Generation using VQ-VAE and Transformers 3 1 20 Apr 2021 not harvested
Latent Video Transformer 1 2 18 Jun 2020 ran 1 of 5 samples (4 unverified)
Transformation-based Adversarial Video Prediction on Large-Scale Data 0 1 9 Mar 2020 not harvested
Exploring Spatial-Temporal Multi-Frequency Analysis for High-Fidelity and Temporal-Consistency Video Prediction 1 1 23 Feb 2020 not harvested
Stochastic Latent Residual Video Prediction 1 1 21 Feb 2020 ran 2 of 12 samples (10 unverified)
Adversarial Video Generation on Complex Datasets 1 2 15 Jul 2019 not harvested
Scaling Autoregressive Video Models 1 1 6 Jun 2019 not harvested
Improved Conditional VRNNs for Video Prediction 1 2 27 Apr 2019 not harvested
VideoFlow: A Conditional Flow-Based Model for Stochastic Video Generation 1 1 4 Mar 2019 not harvested
Stochastic Adversarial Video Prediction 4 4 4 Apr 2018 ran 4 of 16 samples (12 unverified; 3 pointer-only for licence)
Stochastic Video Generation with a Learned Prior 3 3 21 Feb 2018 ran 1 of 2 samples (1 unverified; 2 pointer-only for licence)
Stochastic Variational Video Prediction 3 2 30 Oct 2017 not harvested
MoCoGAN: Decomposing Motion and Content for Video Generation 5 1 17 Jul 2017 not harvested
Unsupervised Learning for Physical Interaction through Video Prediction 2 1 23 May 2016 ran 2 of 2 samples (0 unverified; 2 pointer-only for licence)

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

No modality tagged.

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • BAIR Robot Pushing

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections