Datasets › Something-Something V2

Something-Something V2

Introduced by Raghav Goyal et al. in The "something something" video database for learning and evaluating visual common sense1 Jan 2017 archive 2025-07-28

The 20BN-SOMETHING-SOMETHING V2 dataset is a large collection of labeled video clips that show humans performing pre-defined basic actions with everyday objects. The dataset was created by a large number of crowd workers. It allows machine learning models to develop fine-grained understanding of basic actions that occur in the physical world. It contains 220,847 videos, with 168,913 in the training set, 24,777 in the validation set and 27,157 in the test set. There are 174 labels.

Source

Image Source

Benchmarks archive 2025-07-28

All 7 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Action Recognition Something-Something V2 MVD (Kinetics400 pretrain, ViT-H, 16 frame) Top-1 Accuracy 77.3 Masked Video Distillation: Rethinking Masked Feature... ruiwang2021/mvd +3 123 Compare
Action Recognition In Videos Something-Something V2 STM (16 frames, ImageNet pretraining) Top-1 Accuracy 64.2 STM: SpatioTemporal and Motion Encoding for Action Recognition — 4 Compare
General Action Video Anomaly Detection Something-Something V2 Pooled Image Level kNN Avg. ROC-AUC 0.58 Approaches Toward Physical and General Video Anomaly Detection laurarkart/Physical-Anomalous-Trajectory-or-Motion-PHANTOM-Dataset 3 Compare
Action Classification Something-Something V2 AdaMAE Acc@1 70.04 AdaMAE: Adaptive Masking for Efficient Spatiotemporal... wgcban/adamae +1 1 Compare
Text-to-Video Generation Something-Something V2 MAGVIT FVD 79.1 MAGVIT: Masked Generative Video Transformer google-research/magvit 1 Compare
Video Classification Something-Something V2 MSNet-R50En (ours) Top-5 Accuracy 91 MotionSqueeze: Neural Motion Feature Learning for Video... arunos728/MotionSqueeze +1 1 Compare
Video Prediction Something-Something V2 MAGVIT FVD 28.5 MAGVIT: Masked Generative Video Transformer google-research/magvit 1 Compare

Papers archive 2025-07-28

30 shown of 85 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 290. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
DejaVid: Encoder-Agnostic Learned Temporal Matching for Video Classification 1 1 1 Jan 2025 not harvested
TDS-CLIP: Temporal Difference Side Network for Image-to-Video Transfer Learning 1 1 20 Aug 2024 not harvested
Learning Correlation Structures for Vision Transformers 0 1 5 Apr 2024 not harvested
InternVideo2: Scaling Foundation Models for Multimodal Video Understanding 2 2 22 Mar 2024 not harvested
CAST: Cross-Attention in Space and Time for Video Action Recognition 1 1 30 Nov 2023 ran 9 of 17 samples (8 unverified; 17 pointer-only for licence)
Side4Video: Spatial-Temporal Side Network for Memory-Efficient Image-to-Video Transfer Learning 2 1 27 Nov 2023 not harvested
Asymmetric Masked Distillation for Pre-Training Small Foundation Models 0 2 6 Nov 2023 not harvested
ZeroI2V: Zero-Cost Adaptation of Pre-trained Transformers from Image to Video 2 1 2 Oct 2023 ran 1 of 3 samples (2 unverified)
Temporally-Adaptive Models for Efficient Video Understanding 1 2 10 Aug 2023 ran 2 of 2 samples (0 unverified)
What Can Simple Arithmetic Operations Do for Temporal Modeling? 2 1 18 Jul 2023 ran 3 of 6 samples (3 unverified)
Hiera: A Hierarchical Vision Transformer without the Bells-and-Whistles 4 1 1 Jun 2023 ran 0 of 6 samples (6 unverified)
SCP: Soft Conditional Prompt Learning for Aerial Video Action Recognition 0 1 21 May 2023 not harvested
Implicit Temporal Modeling with Learnable Alignment for Video Recognition 1 2 20 Apr 2023 ran 2 of 4 samples (2 unverified; 4 pointer-only for licence)
VideoMAE V2: Scaling Video Masked Autoencoders with Dual Masking 1 1 29 Mar 2023 ran 2 of 6 samples (4 unverified)
The effectiveness of MAE pre-pretraining for billion-scale pretraining 1 1 23 Mar 2023 not harvested
ViC-MAE: Self-Supervised Representation Learning from Images and Video with Contrastive Masked Autoencoders 2 1 21 Mar 2023 not harvested
Maximizing Spatio-Temporal Entropy of Deep 3D CNNs for Efficient Video Recognition 1 1 5 Mar 2023 not harvested
Multi-scale Motion-Aware Module for Video Action Recognition 0 1 19 Feb 2023 not harvested
MAGVIT: Masked Generative Video Transformer 1 2 10 Dec 2022 ran 1 of 10 samples (9 unverified)
Masked Video Distillation: Rethinking Masked Feature Modeling for Self-supervised Video Representation Learning 4 4 8 Dec 2022 not harvested
Rethinking Video ViTs: Sparse Video Tubes for Joint Image and Video Learning 1 1 6 Dec 2022 ran 1 of 3 samples (2 unverified)
InternVideo: General Video Foundation Models via Generative and Discriminative Learning 2 1 6 Dec 2022 ran 3 of 3 samples (0 unverified)
Global Temporal Difference Network for Action Recognition 0 1 23 Nov 2022 not harvested
AdaMAE: Adaptive Masking for Efficient Spatiotemporal Learning with Masked Autoencoders 2 1 16 Nov 2022 ran 9 of 25 samples (16 unverified)
UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer 2 1 22 Sep 2022 not harvested
OmniVL:One Foundation Model for Image-Language and Video-Language Tasks 0 1 15 Sep 2022 not harvested
Spatial-Temporal Pyramid Graph Reasoning for Action Recognition 0 1 9 Aug 2022 not harvested
Spatiotemporal Self-attention Modeling with Temporal Patch Shift for Action Recognition 1 1 27 Jul 2022 ran 1 of 1 samples (0 unverified)
MAR: Masked Autoencoders for Efficient Action Recognition 1 4 24 Jul 2022 not harvested
Action Recognition With Motion Diversification and Dynamic Selection 0 1 15 Jul 2022 not harvested

The full list of 85 is in the JSON twin.

Dataset loaders archive 2025-07-28

1 loader as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

Custom

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • Something-Something V2

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections