Datasets › Kinetics-600

Kinetics-600

Introduced by Joao Carreira et al. in A Short Note about Kinetics-6001 Jan 2018 archive 2025-07-28

The Kinetics-600 is a large-scale action recognition dataset which consists of around 480K videos from 600 action categories. The 480K videos are divided into 390K, 30K, 60K for training, validation and test sets, respectively. Each video in the dataset is a 10-second clip of action moment annotated from raw YouTube video. It is an extensions of the Kinetics-400 dataset.

Source: Learning to Localize Actions from Moments Image Source: https://towardsdatascience.com/downloading-the-kinetics-dataset-for-human-action-recognition-in-deep-learning-500c3d50f776

Benchmarks archive 2025-07-28

All 3 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Action Classification Kinetics-600 InternVideo2-6B Top-1 Accuracy 91.9 InternVideo2: Scaling Foundation Models for Multimodal... opengvlab/internvideo +1 65 Compare
Self-Supervised Action Recognition Kinetics-600 CVRL (R3D-152 2x) Top-1 Accuracy 72.9 Spatiotemporal Contrastive Video Representation Learning tensorflow/models +3 5 Compare
Action Recognition In Videos Kinetics-600 Florence Top-1 Accuracy 87.8 Florence: A New Foundation Model for Computer Vision microsoft/unicl +1 1 Compare

Papers archive 2025-07-28

30 shown of 35 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 148. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
InternVideo2: Scaling Foundation Models for Multimodal Video Understanding 2 2 22 Mar 2024 not harvested
Hiera: A Hierarchical Vision Transformer without the Bells-and-Whistles 4 1 1 Jun 2023 ran 0 of 6 samples (6 unverified)
VideoMAE V2: Scaling Video Masked Autoencoders with Dual Masking 1 2 29 Mar 2023 ran 2 of 6 samples (4 unverified)
Unmasked Teacher: Towards Training-Efficient Video Foundation Models 1 1 28 Mar 2023 ran 3 of 8 samples (5 unverified)
mPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and Video 4 1 1 Feb 2023 ran 9 of 19 samples (10 unverified)
Rethinking Video ViTs: Sparse Video Tubes for Joint Image and Video Learning 1 3 6 Dec 2022 ran 1 of 3 samples (2 unverified)
InternVideo: General Video Foundation Models via Generative and Discriminative Learning 2 1 6 Dec 2022 ran 3 of 3 samples (0 unverified)
EVA: Exploring the Limits of Masked Visual Representation Learning at Scale 6 1 14 Nov 2022 ran 1 of 3 samples (2 unverified)
UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer 2 1 22 Sep 2022 not harvested
Expanding Language-Image Pretrained Models for General Video Recognition 2 1 4 Aug 2022 ran 3 of 3 samples (0 unverified; 3 pointer-only for licence)
CoCa: Contrastive Captioners are Image-Text Foundation Models 6 2 4 May 2022 ran 9 of 17 samples (8 unverified)
Multiview Transformers for Video Recognition 1 1 12 Jan 2022 not harvested
MERLOT Reserve: Neural Script Knowledge through Vision and Language and Sound 0 4 7 Jan 2022 not harvested
Masked Feature Prediction for Self-Supervised Visual Pre-Training 6 1 16 Dec 2021 not harvested
Co-training Transformer with Videos and Images Improves Action Recognition 0 2 14 Dec 2021 not harvested
MViTv2: Improved Multiscale Vision Transformers for Classification and Detection 9 4 2 Dec 2021 not harvested
Florence: A New Foundation Model for Computer Vision 2 2 22 Nov 2021 not harvested
UniFormer: Unified Transformer for Efficient Spatial-Temporal Representation Learning 3 1 29 Sep 2021 not harvested
Revisiting 3D ResNets for Video Recognition 5 1 3 Sep 2021 not harvested
Video Swin Transformer 15 2 24 Jun 2021 ran 7 of 32 samples (25 unverified)
TokenLearner: What Can 8 Learned Tokens Do for Images and Videos? 11 1 21 Jun 2021 ran 3 of 3 samples (0 unverified)
Space-time Mixing Attention for Video Transformer 1 1 10 Jun 2021 ran 1 of 1 samples (0 unverified)
VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and Text 5 1 22 Apr 2021 ran 5 of 8 samples (3 unverified; 8 pointer-only for licence)
Multiscale Vision Transformers 8 3 22 Apr 2021 ran 13 of 26 samples (13 unverified; 5 pointer-only for licence)
Broaden Your Views for Self-Supervised Video Learning 1 1 30 Mar 2021 ran 0 of 8 samples (8 unverified)
ViViT: A Video Vision Transformer 10 3 29 Mar 2021 ran 14 of 21 samples (7 unverified; 1 pointer-only for licence)
MoViNets: Mobile Video Networks for Efficient Video Recognition 3 8 21 Mar 2021 ran 8 of 13 samples (5 unverified)
PERF-Net: Pose Empowered RGB-Flow Net 0 1 28 Sep 2020 not harvested
Spatiotemporal Contrastive Video Representation Learning 4 3 9 Aug 2020 not harvested
Self-Supervised MultiModal Versatile Networks 1 1 29 Jun 2020 not harvested

The full list of 35 is in the JSON twin.

Dataset loaders archive 2025-07-28

2 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

CC BY 4.0

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • Kinetics-600 12 frames, 64x64
  • Kinetics-600 48 frames, 64x64
  • Kinetics-600 12 frames, 128x128
  • Kinetics-600

4 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections