Datasets › Something-Something V1

Something-Something V1

Introduced by Raghav Goyal et al. in The "something something" video database for learning and evaluating visual common sense1 Jan 2017 archive 2025-07-28

The 20BN-SOMETHING-SOMETHING dataset is a large collection of labeled video clips that show humans performing pre-defined basic actions with everyday objects. The dataset was created by a large number of crowd workers. It allows machine learning models to develop fine-grained understanding of basic actions that occur in the physical world. It contains 108,499 videos, with 86,017 in the training set, 11,522 in the validation set and 10,960 in the test set. There are 174 labels.

⚠️ Attention: This is the outdated V1 of the dataset. V2 is available here.

Source: https://20bn.com/datasets/something-something/v1 Image Source: https://20bn.com/datasets/something-something/v1

Benchmarks archive 2025-07-28

All 3 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Action Recognition Something-Something V1 InternVideo Top 1 Accuracy 70.0 InternVideo: General Video Foundation Models via... opengvlab/internvideo +1 74 Compare
Action Recognition In Videos Something-Something V1 STM (16 frames, ImageNet pretraining) Top 1 Accuracy 50.7 STM: SpatioTemporal and Motion Encoding for Action Recognition — 3 Compare
Video Classification Something-Something V1 MSNet-R50En (ours) Top-5 Accuracy 84 MotionSqueeze: Neural Motion Feature Learning for Video... arunos728/MotionSqueeze +1 1 Compare

Papers archive 2025-07-28

30 shown of 46 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 117. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
TDS-CLIP: Temporal Difference Side Network for Image-to-Video Transfer Learning 1 1 20 Aug 2024 not harvested
Learning Correlation Structures for Vision Transformers 0 1 5 Apr 2024 not harvested
Side4Video: Spatial-Temporal Side Network for Memory-Efficient Image-to-Video Transfer Learning 2 1 27 Nov 2023 not harvested
Temporally-Adaptive Models for Efficient Video Understanding 1 2 10 Aug 2023 ran 2 of 2 samples (0 unverified)
What Can Simple Arithmetic Operations Do for Temporal Modeling? 2 1 18 Jul 2023 ran 3 of 6 samples (3 unverified)
VideoMAE V2: Scaling Video Masked Autoencoders with Dual Masking 1 1 29 Mar 2023 ran 2 of 6 samples (4 unverified)
Multi-scale Motion-Aware Module for Video Action Recognition 0 1 19 Feb 2023 not harvested
InternVideo: General Video Foundation Models via Generative and Discriminative Learning 2 1 6 Dec 2022 ran 3 of 3 samples (0 unverified)
UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer 2 1 22 Sep 2022 not harvested
Spatial-Temporal Pyramid Graph Reasoning for Action Recognition 0 1 9 Aug 2022 not harvested
Spatiotemporal Self-attention Modeling with Temporal Patch Shift for Action Recognition 1 1 27 Jul 2022 ran 1 of 1 samples (0 unverified)
AE-Net:Adjoint Enhancement Network for Efficient Action Recognition in Video Understanding 0 1 21 Jul 2022 not harvested
Action Recognition With Motion Diversification and Dynamic Selection 0 1 15 Jul 2022 not harvested
Stand-Alone Inter-Frame Attention in Video Models 1 1 14 Jun 2022 not harvested
MLP-3D: A MLP-like 3D Architecture with Grouped Time Mixing 0 1 13 Jun 2022 not harvested
Motion-driven Visual Tempo Learning for Video-based Action Recognition 2 1 24 Feb 2022 not harvested
Action Keypoint Network for Efficient Video Recognition 0 1 17 Jan 2022 not harvested
Relational Self-Attention: What's Missing in Attention for Video Understanding 1 4 2 Nov 2021 ran 1 of 1 samples (0 unverified; 1 pointer-only for licence)
UniFormer: Unified Transformer for Efficient Spatial-Temporal Representation Learning 3 2 29 Sep 2021 not harvested
EAN: Event Adaptive Network for Enhanced Action Recognition 1 1 22 Jul 2021 not harvested
CT-Net: Channel Tensorization Network for Video Classification 1 1 3 Jun 2021 ran 10 of 13 samples (3 unverified)
Busy-Quiet Video Disentangling for Video Classification 2 1 29 Mar 2021 not harvested
Learning Self-Similarity in Space and Time as Generalized Motion for Video Action Recognition 1 3 14 Feb 2021 ran 3 of 4 samples (1 unverified; 4 pointer-only for licence)
TDN: Temporal Difference Networks for Efficient Action Recognition 1 1 18 Dec 2020 ran 2 of 4 samples (2 unverified)
MVFNet: Multi-View Fusion Network for Efficient Video Recognition 3 1 13 Dec 2020 not harvested
Diverse Temporal Aggregation and Depthwise Spatiotemporal Factorization for Efficient Video Classification 1 6 1 Dec 2020 not harvested
PAN: Towards Fast Action Recognition via Learning Persistence of Appearance 2 1 8 Aug 2020 not harvested
MotionSqueeze: Neural Motion Feature Learning for Video Understanding 2 5 20 Jul 2020 ran 3 of 8 samples (5 unverified)
Region-based Non-local Operation for Video Classification 1 2 17 Jul 2020 not harvested
Knowing What, Where and When to Look: Efficient Video Action Modeling with Attention 0 1 2 Apr 2020 not harvested

The full list of 46 is in the JSON twin.

Dataset loaders archive 2025-07-28

1 loader as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • Something-Something V1

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections