Datasets › Kinetics-700

Kinetics-700

Introduced by Joao Carreira et al. in A Short Note on the Kinetics-700 Human Action Dataset15 Jul 2019 archive 2025-07-28

Kinetics-700 is a video dataset of 650,000 clips that covers 700 human action classes. The videos include human-object interactions such as playing instruments, as well as human-human interactions such as shaking hands and hugging. Each action class has at least 700 video clips. Each clip is annotated with an action class and lasts approximately 10 seconds.

Benchmarks archive 2025-07-28

All 3 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Action Classification Kinetics-700 InternVideo2-6B Top-1 Accuracy 85.9 InternVideo2: Scaling Foundation Models for Multimodal... opengvlab/internvideo +1 36 Compare
Image Clustering Kinetics-700 TURTLE (CLIP + DINOv2) Accuracy 43.0 Let Go of Your Labels with Unsupervised Transfer mlbio-epfl/turtle 1 Compare
Video Generation Kinetics-700 DiT-XL/2 + CVAE-FT-SE FID 8.59 Improving the Diffusability of Autoencoders — 1 Compare

Papers archive 2025-07-28

20 shown of 20 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 95. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Improving the Diffusability of Autoencoders 0 1 20 Feb 2025 not harvested
Let Go of Your Labels with Unsupervised Transfer 1 1 11 Jun 2024 ran 3 of 4 samples (1 unverified; 4 pointer-only for licence)
InternVideo2: Scaling Foundation Models for Multimodal Video Understanding 2 2 22 Mar 2024 not harvested
Hiera: A Hierarchical Vision Transformer without the Bells-and-Whistles 4 1 1 Jun 2023 ran 0 of 6 samples (6 unverified)
Unmasked Teacher: Towards Training-Efficient Video Foundation Models 1 1 28 Mar 2023 ran 3 of 8 samples (5 unverified)
AIM: Adapting Image Models for Efficient Video Action Recognition 1 1 6 Feb 2023 not harvested
mPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and Video 4 1 1 Feb 2023 ran 9 of 19 samples (10 unverified)
Rethinking Video ViTs: Sparse Video Tubes for Joint Image and Video Learning 1 1 6 Dec 2022 ran 1 of 3 samples (2 unverified)
InternVideo: General Video Foundation Models via Generative and Discriminative Learning 2 1 6 Dec 2022 ran 3 of 3 samples (0 unverified)
EVA: Exploring the Limits of Masked Visual Representation Learning at Scale 6 1 14 Nov 2022 ran 1 of 3 samples (2 unverified)
UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer 2 1 22 Sep 2022 not harvested
CoCa: Contrastive Captioners are Image-Text Foundation Models 6 2 4 May 2022 ran 9 of 17 samples (8 unverified)
Vision Models Are More Robust And Fair When Pretrained On Uncurated Images Without Supervision 1 1 16 Feb 2022 not harvested
Multiview Transformers for Video Recognition 1 1 12 Jan 2022 not harvested
Masked Feature Prediction for Self-Supervised Visual Pre-Training 6 1 16 Dec 2021 not harvested
Co-training Transformer with Videos and Images Improves Action Recognition 0 2 14 Dec 2021 not harvested
MViTv2: Improved Multiscale Vision Transformers for Classification and Detection 9 3 2 Dec 2021 not harvested
VidTr: Video Transformer Without Convolutions 0 4 23 Apr 2021 not harvested
MoViNets: Mobile Video Networks for Efficient Video Recognition 3 7 21 Mar 2021 ran 8 of 13 samples (5 unverified)
Learn to cycle: Time-consistent feature discovery for action recognition 1 5 15 Jun 2020 not harvested

Dataset loaders archive 2025-07-28

2 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

Commons Attribution 4.0 International License

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • Kinetics-700

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections