Papers › The "something something" video database for learning and evaluating visual common sense

The "something something" video database for learning and evaluating visual common sense

13 Jun 2017ICCV 2017 10arXiv:1706.04261archive 2025-07-28

Raghav Goyal, Samira Ebrahimi Kahou, Vincent Michalski, Joanna Materzyńska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ingo Fruend, Peter Yianilos, Moritz Mueller-Freitag, Florian Hoppe, Christian Thurau, Ingo Bax, Roland Memisevic

Neural networks trained on datasets such as ImageNet have led to major advances in visual object classification. One obstacle that prevents networks from reasoning more deeply about complex scenes and situations, and from integrating visual knowledge with natural language, like humans do, is their lack of common sense knowledge about the physical world. Videos, unlike still images, contain a wealth of detailed information about the physical world. However, most labelled video datasets represent high-level concepts rather than detailed physical aspects about actions and scenes. In this work, we describe our ongoing collection of the "something-something" database of video prediction tasks whose solutions require a common sense understanding of the depicted situation. The database currently contains more than 100,000 videos across 174 classes, which are defined as caption-templates. We also describe the challenges in crowd-sourcing this data at scale.

PaperPDFConference PDFCode

In Syntology View this paper on Syntology: its repositories, every harvested function with whether it ran, its licence and the call to fetch it.

Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

akshyta/Human-Activity-Recognition mentioned on GitHubpytorch report
bit-ml/dyreg-gnn mentioned on GitHubpytorch report
caspillaga/Conv3DSelfAttention mentioned on GitHubpytorch report
jayleicn/singularity mentioned on GitHubpytorch report
latte488/smth-smth-v2 mentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Action RecognitionCommon Sense ReasoningGeneral ClassificationVideo Prediction

Datasets

Introduced by this paper, per the archive.

Something-Something V1Something-Something V2

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Action Recognition Something-Something V2 model3D_1 with left-right augmentation and fps jitter Top-1 Accuracy 51.33 #116 of 123 Archive leaderboard report
Action Recognition Something-Something V2 model3D_1 with left-right augmentation and fps jitter Top-5 Accuracy 80.46 #116 of 123 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections