Datasets › EPIC-KITCHENS-100
EPIC-KITCHENS-100
This paper introduces the pipeline to scale the largest dataset in egocentric vision EPIC-KITCHENS. The effort culminates in EPIC-KITCHENS-100, a collection of 100 hours, 20M frames, 90K actions in 700 variable-length videos, capturing long-term unscripted activities in 45 environments, using head-mounted cameras. Compared to its previous version (EPIC-KITCHENS-55), EPIC-KITCHENS-100 has been annotated using a novel pipeline that allows denser (54% more actions per minute) and more complete annotations of fine-grained actions (+128% more action segments). This collection also enables evaluating the "test of time" - i.e. whether models trained on data collected in 2018 can generalise to new footage collected under the same hypotheses albeit "two years on". The dataset is aligned with 6 challenges: action recognition (full and weak supervision), action detection, action anticipation, cross-modal retrieval (from captions), as well as unsupervised domain adaptation for action recognition. For each challenge, we define the task, provide baselines and evaluation metrics.
Benchmarks archive 2025-07-28
All 6 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.
| First row (archive order) | Paper | Code | ||||
|---|---|---|---|---|---|---|
| Action Recognition | EPIC-KITCHENS-100 | LLaVAction Action@1 58.3 | LLaVAction: evaluating and training multi-modal large... | adaptivemotorcontrollab/llavaction | 32 | Compare |
| Action Anticipation | EPIC-KITCHENS-100 | PlausiVL Recall@5 27.60 | Can't make an Omelette without Breaking some Eggs:... | — | 9 | Compare |
| Temporal Action Localization | EPIC-KITCHENS-100 | AdaTAD (verb, VideoMAE-L) Avg mAP (0.1-0.5) 29.3 | End-to-End Temporal Action Detection with 1B Parameters... | sming256/OpenTAD +1 | 6 | Compare |
| Unsupervised Domain Adaptation | EPIC-KITCHENS-100 | TranSVAE Average Accuracy 52.6 | — | — | 5 | Compare |
| Audio Classification | EPIC-KITCHENS-100 | Audiovisual Masked Autoencoder (Audiovisual, Single) Top-1 Action 46.0 | Audiovisual Masked Autoencoders | google-research/scenic +1 | 4 | Compare |
| Open Vocabulary Action Recognition | EPIC-KITCHENS-100 | OAP+AOP HM 17.0 | Opening the Vocabulary of Egocentric Actions | dibschat/openvocab-egoAR | 1 | Compare |
Papers archive 2025-07-28
30 shown of 40 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 162. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.
The full list of 40 is in the JSON twin.
Dataset loaders archive 2025-07-28
No loader listed in the archive.
Tasks archive 2025-07-28
License archive 2025-07-28
Modalities archive 2025-07-28
Languages archive 2025-07-28
Variants archive 2025-07-28
- EPIC-KITCHENS-100
1 variant name, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections