Datasets › Kinetics › Papers, page 2

Kinetics (Kinetics Human Action Video Dataset)

Papers archive 2025-07-28

papers with a benchmark row: 171 · with a code link: 132 · where Syntology ran a sample: 73 (66 with a run with no instrument failure, 7 where every run was a failure of Syntology's instrument) Syntology

Show: all papers with a benchmark rowonly where code ran (73 of 171 with a benchmark row: 66 with a run with no instrument failure, 7 where every run was a failure of Syntology's instrument)

Page 2 of 2: papers 101 to 171 of 171 with a leaderboard row on this dataset's benchmarks, newest first by the archive's date (ties by slug; undated papers last).

The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset, not that list; the archive's count for this dataset is 1,341. The Syntology column is from Syntology's graph, stated per sample; it is not part of any archive number. A line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified” (C is a part of N, never taken away from it); the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. When the archive marks a repository official for the paper, the cell starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record.

PaperCodeResultsDateSamples run Syntology
VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and Text 5 1 22 Apr 2021 community repositories only · 5 ran (of which 4 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (8 pointer-only for licence)
ViViT: A Video Vision Transformer 10 2 29 Mar 2021 official: no sample here; runs from other or unrecorded repositories · 14 ran (of which 10 constructed an object rather than computing a result; 13 with no instrument failure: 1 honoured, 0 violated, 12 with no contract checked; 1 where Syntology's instrument failed) · 7 unverified (1 pointer-only for licence)
Busy-Quiet Video Disentangling for Video Classification 2 1 29 Mar 2021 not harvested
An Image is Worth 16x16 Words, What is a Video Worth? 2 2 25 Mar 2021 official: harvested, nothing ran · 0 ran · 1 unverified
MoViNets: Mobile Video Networks for Efficient Video Recognition 3 7 21 Mar 2021 community repositories only · 9 ran (of which 6 constructed an object rather than computing a result; 8 with no instrument failure: 1 honoured, 0 violated, 7 with no contract checked; 1 where Syntology's instrument failed) · 4 unverified
Predicting Video with VQVAE 1 1 2 Mar 2021 not harvested
Is Space-Time Attention All You Need for Video Understanding? 16 3 9 Feb 2021 official: no sample here; runs from other or unrecorded repositories · 35 ran (of which 17 constructed an object rather than computing a result; 21 with no instrument failure: 0 honoured, 1 violated, 20 with no contract checked; 14 where Syntology's instrument failed) · 8 unverified (14 pointer-only for licence)
Video Transformer Network 1 4 1 Feb 2021 not harvested
TDN: Temporal Difference Networks for Efficient Action Recognition 1 1 18 Dec 2020 official (archive's flag): 2 ran · 2 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified
MVFNet: Multi-View Fusion Network for Efficient Video Recognition 3 1 13 Dec 2020 not harvested
Pose Refinement Graph Convolutional Network for Skeleton-based Action Recognition 0 1 14 Oct 2020 not harvested
All About Knowledge Graphs for Actions 0 1 28 Aug 2020 not harvested
Skeleton-based Action Recognition via Spatial and Temporal Transformer Networks 1 1 17 Aug 2020 not harvested
Spatiotemporal Contrastive Video Representation Learning 4 3 9 Aug 2020 not harvested
Dynamic GCN: Context-enriched Topology Learning for Skeleton-based Action Recognition 1 1 29 Jul 2020 not harvested
MotionSqueeze: Neural Motion Feature Learning for Video Understanding 2 1 20 Jul 2020 6 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (2 pointer-only for licence)
Region-based Non-local Operation for Video Classification 1 1 17 Jul 2020 not harvested
Latent Video Transformer 1 1 18 Jun 2020 official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (1 pointer-only for licence)
X3D: Expanding Architectures for Efficient Video Recognition 8 4 9 Apr 2020 community repositories only · 11 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 1 honoured, 0 violated, 9 with no contract checked; 1 where Syntology's instrument failed) · 4 unverified
Disentangling and Unifying Graph Convolutions for Skeleton-Based Action Recognition 3 1 31 Mar 2020 official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified
Omni-sourced Webly-supervised Learning for Video Recognition 3 3 29 Mar 2020 official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified
Temporal Extension Module for Skeleton-Based Action Recognition 0 1 19 Mar 2020 not harvested
Predictively Encoded Graph Convolutional Network for Noise-Robust Skeleton-based Action Recognition 1 1 17 Mar 2020 not harvested
Transformation-based Adversarial Video Prediction on Large-Scale Data 0 1 9 Mar 2020 not harvested
Unifying Graph Embedding Features with Graph Convolutional Networks for Skeleton-based Action Recognition 0 1 6 Mar 2020 not harvested
Skeleton-Based Action Recognition with Multi-Stream Adaptive Graph Convolutional Networks 2 2 15 Dec 2019 not harvested
A Multigrid Method for Efficiently Training Video Models 3 1 2 Dec 2019 community repositories only · 6 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 1 honoured, 0 violated, 4 with no contract checked; 1 where Syntology's instrument failed) · 4 unverified (3 pointer-only for licence)
More Is Less: Learning Efficient Video Representations by Big-Little Network and Depthwise Temporal Aggregation 1 1 2 Dec 2019 official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified
An Attention-Enhanced Recurrent Graph Convolutional Network for Skeleton-Based Action Recognition 0 1 27 Nov 2019 not harvested
Learning Graph Convolutional Network for Skeleton-based Human Action Recognition by Neural Searching 1 1 11 Nov 2019 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (2 pointer-only for licence)
Chirality Nets for Human Pose Regression 1 1 31 Oct 2019 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified
Action recognition with spatial-temporal discriminative filter banks 0 1 20 Aug 2019 not harvested
STM: SpatioTemporal and Motion Encoding for Action Recognition 0 1 7 Aug 2019 not harvested
Two-Stream Video Classification with Cross-Modality Attention 0 1 1 Aug 2019 not harvested
Spatiotemporal graph routing for skeleton-based action recognition 0 1 17 Jul 2019 not harvested
Adversarial Video Generation on Complex Datasets 1 4 15 Jul 2019 not harvested
Learning Spatio-Temporal Representation with Local and Global Diffusion 0 3 13 Jun 2019 not harvested
FASTER Recurrent Networks for Efficient Video Classification 0 2 10 Jun 2019 not harvested
Video Modeling with Correlation Networks 0 1 7 Jun 2019 not harvested
Scaling Autoregressive Video Models 1 1 6 Jun 2019 not harvested
Global Textual Relation Embedding for Relational Understanding 1 1 3 Jun 2019 not harvested
Collaborative Spatiotemporal Feature Learning for Video Action Recognition 1 1 1 Jun 2019 not harvested
MARS: Motion-Augmented RGB Stream for Action Recognition 1 3 1 Jun 2019 not harvested
Skeleton-Based Action Recognition With Directed Graph Neural Networks 1 1 1 Jun 2019 not harvested
What Makes Training Multi-Modal Classification Networks Hard? 3 2 29 May 2019 5 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 1 where Syntology's instrument failed) · 6 unverified
Large-scale weakly-supervised pre-training for video action recognition 3 1 2 May 2019 12 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 0 violated, 12 with no contract checked; 0 where Syntology's instrument failed) · 9 unverified
Actional-Structural Graph Convolutional Networks for Skeleton-based Action Recognition 1 1 26 Apr 2019 official (archive's flag): 14 ran · 14 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 2 honoured, 1 violated, 9 with no contract checked; 2 where Syntology's instrument failed) · 4 unverified (5 pointer-only for licence)
Large Scale Holistic Video Understanding 1 1 25 Apr 2019 not harvested
Drop an Octave: Reducing Spatial Redundancy in Convolutional Neural Networks with Octave Convolution 28 1 10 Apr 2019 community repositories only · 24 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 1 honoured, 0 violated, 10 with no contract checked; 13 where Syntology's instrument failed) · 10 unverified (9 pointer-only for licence)
Video Classification with Channel-Separated Convolutional Networks 7 5 4 Apr 2019 official (archive's flag): 1 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified
D3D: Distilled 3D Networks for Video Action Recognition 1 2 19 Dec 2018 not harvested
SlowFast Networks for Video Recognition 15 6 10 Dec 2018 community repositories only · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (1 pointer-only for licence)
Evolving Space-Time Neural Architectures for Videos 0 1 26 Nov 2018 not harvested
TSM: Temporal Shift Module for Efficient Video Understanding 13 1 20 Nov 2018 official: harvested, nothing ran · 10 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 1 honoured, 0 violated, 4 with no contract checked; 5 where Syntology's instrument failed) · 6 unverified (6 pointer-only for licence)
Skeleton-Based Action Recognition with Synchronous Local and Non-local Spatio-temporal Learning and Frequency Attention 0 1 10 Nov 2018 not harvested
A²-Nets: Double Attention Networks 0 1 27 Oct 2018 not harvested
Representation Flow for Action Recognition 5 1 2 Oct 2018 not harvested
Multi-Fiber Networks for Video Recognition 0 1 30 Jul 2018 not harvested
Rethinking Spatiotemporal Feature Learning: Speed-Accuracy Trade-offs in Video Classification 2 3 13 Dec 2017 not harvested
A Closer Look at Spatiotemporal Convolutions for Action Recognition 24 6 30 Nov 2017 community repositories only · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (4 pointer-only for licence)
Appearance-and-Relation Networks for Video Classification 1 1 24 Nov 2017 not harvested
Non-local Neural Networks 32 1 21 Nov 2017 official (archive's flag): 1 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (4 pointer-only for licence)
ConvNet Architecture Search for Spatiotemporal Feature Learning 1 1 16 Aug 2017 not harvested
Revisiting the Effectiveness of Off-the-shelf Temporal Modeling Approaches for Large-scale Video Classification 0 1 12 Aug 2017 not harvested
Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset 34 1 22 May 2017 16 ran (of which 7 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 6 where Syntology's instrument failed) · 12 unverified (7 pointer-only for licence)
Learning a Deep Embedding Model for Zero-Shot Learning 4 1 15 Nov 2016 not harvested
Temporal Segment Networks: Towards Good Practices for Deep Action Recognition 22 1 2 Aug 2016 community repositories only · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 2 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 14 unverified (3 pointer-only for licence)
An embarrassingly simple approach to zero-shot learning 1 1 6 Jul 2015 not harvested
Label-Embedding for Image Classification 2 1 30 Mar 2015 not harvested
Evaluation of Output Embeddings for Fine-Grained Image Classification 2 1 30 Sep 2014 not harvested
DeViSE: A Deep Visual-Semantic Embedding Model 0 1 1 Dec 2013 not harvested