Datasets › EPIC-KITCHENS-100

EPIC-KITCHENS-100

Introduced by Dima Damen et al. in Rescaling Egocentric Vision23 Jun 2020 archive 2025-07-28

This paper introduces the pipeline to scale the largest dataset in egocentric vision EPIC-KITCHENS. The effort culminates in EPIC-KITCHENS-100, a collection of 100 hours, 20M frames, 90K actions in 700 variable-length videos, capturing long-term unscripted activities in 45 environments, using head-mounted cameras. Compared to its previous version (EPIC-KITCHENS-55), EPIC-KITCHENS-100 has been annotated using a novel pipeline that allows denser (54% more actions per minute) and more complete annotations of fine-grained actions (+128% more action segments). This collection also enables evaluating the "test of time" - i.e. whether models trained on data collected in 2018 can generalise to new footage collected under the same hypotheses albeit "two years on". The dataset is aligned with 6 challenges: action recognition (full and weak supervision), action detection, action anticipation, cross-modal retrieval (from captions), as well as unsupervised domain adaptation for action recognition. For each challenge, we define the task, provide baselines and evaluation metrics.

Benchmarks archive 2025-07-28

All 6 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Action Recognition EPIC-KITCHENS-100 LLaVAction Action@1 58.3 LLaVAction: evaluating and training multi-modal large... adaptivemotorcontrollab/llavaction 32 Compare
Action Anticipation EPIC-KITCHENS-100 PlausiVL Recall@5 27.60 Can't make an Omelette without Breaking some Eggs:... — 9 Compare
Temporal Action Localization EPIC-KITCHENS-100 AdaTAD (verb, VideoMAE-L) Avg mAP (0.1-0.5) 29.3 End-to-End Temporal Action Detection with 1B Parameters... sming256/OpenTAD +1 6 Compare
Unsupervised Domain Adaptation EPIC-KITCHENS-100 TranSVAE Average Accuracy 52.6 — — 5 Compare
Audio Classification EPIC-KITCHENS-100 Audiovisual Masked Autoencoder (Audiovisual, Single) Top-1 Action 46.0 Audiovisual Masked Autoencoders google-research/scenic +1 4 Compare
Open Vocabulary Action Recognition EPIC-KITCHENS-100 OAP+AOP HM 17.0 Opening the Vocabulary of Egocentric Actions dibschat/openvocab-egoAR 1 Compare

Papers archive 2025-07-28

30 shown of 40 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 162. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
LLaVAction: evaluating and training multi-modal large language models for action recognition 1 1 24 Mar 2025 not harvested
Extending Video Masked Autoencoders to 128 frames 0 1 20 Nov 2024 not harvested
Semantically Guided Representation Learning For Action Anticipation 1 1 2 Jul 2024 ran 6 of 8 samples (2 unverified)
Can't make an Omelette without Breaking some Eggs: Plausible Action Anticipation using Large Video-Language Models 0 1 30 May 2024 not harvested
TIM: A Time Interval Machine for Audio-Visual Action Recognition 1 1 8 Apr 2024 ran 11 of 12 samples (1 unverified; 12 pointer-only for licence)
Uncertainty-aware Action Decoupling Transformer for Action Anticipation 0 1 1 Jan 2024 not harvested
CAST: Cross-Attention in Space and Time for Video Action Recognition 1 1 30 Nov 2023 ran 9 of 17 samples (8 unverified; 17 pointer-only for licence)
End-to-End Temporal Action Detection with 1B Parameters Across 1000 Frames 2 1 28 Nov 2023 not harvested
Training a Large Video Model on a Single Machine in a Day 1 1 28 Sep 2023 not harvested
Opening the Vocabulary of Egocentric Actions 1 1 22 Aug 2023 not harvested
Temporally-Adaptive Models for Efficient Video Understanding 1 2 10 Aug 2023 ran 2 of 2 samples (0 unverified)
TemporalMaxer: Maximize Temporal Context with only Max Pooling for Temporal Action Localization 1 1 16 Mar 2023 not harvested
TriDet: Temporal Action Detection with Relative Boundary Modeling 1 1 13 Mar 2023 ran 5 of 15 samples (10 unverified)
Audiovisual Masked Autoencoders 2 3 9 Dec 2022 not harvested
Learning Video Representations from Large Language Models 3 1 8 Dec 2022 ran 6 of 20 samples (14 unverified; 20 pointer-only for licence)
Interaction Region Visual Transformer for Egocentric Action Anticipation 1 1 25 Nov 2022 not harvested
Anticipative Feature Fusion Transformer for Multi-Modal Action Anticipation 1 1 23 Oct 2022 not harvested
Play It Back: Iterative Attention for Audio Recognition 1 1 20 Oct 2022 not harvested
Multiscale Multimodal Transformer for Multimodal Action Recognition 0 1 22 Sep 2022 not harvested
AVT: Audio-Video Transformer for Multimodal Action Recognition 0 1 22 Sep 2022 not harvested
M&M Mix: A Multimodal Multiview Transformer Ensemble 0 1 20 Jun 2022 not harvested
Gate-Shift-Fuse for Video Action Recognition 2 1 16 Mar 2022 not harvested
ActionFormer: Localizing Moments of Actions with Transformers 1 1 16 Feb 2022 not harvested
Omnivore: A Single Model for Many Visual Modalities 2 1 20 Jan 2022 ran 2 of 2 samples (0 unverified; 1 pointer-only for licence)
MeMViT: Memory-Augmented Multiscale Vision Transformer for Efficient Long-Term Video Recognition 1 2 20 Jan 2022 not harvested
Multiview Transformers for Video Recognition 1 1 12 Jan 2022 not harvested
Contrast and Mix: Temporal Contrastive Video Domain Adaptation with Background Mixing 0 1 28 Oct 2021 not harvested
Object-Region Video Transformers 1 1 13 Oct 2021 ran 1 of 7 samples (6 unverified; 7 pointer-only for licence)
Attention Bottlenecks for Multimodal Fusion 1 1 30 Jun 2021 not harvested
Keeping Your Eye on the Ball: Trajectory Attention in Video Transformers 2 3 9 Jun 2021 ran 1 of 2 samples (1 unverified; 2 pointer-only for licence)

The full list of 40 is in the JSON twin.

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

CC BY NC 4.0

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • EPIC-KITCHENS-100

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections