Datasets › YouCook2

YouCook2

Introduced by Luowei Zhou et al. in Towards Automatic Learning of Procedures from Web Instructional Videos1 Jan 2018 archive 2025-07-28

YouCook2 is the largest task-oriented, instructional video dataset in the vision community. It contains 2000 long untrimmed videos from 89 cooking recipes; on average, each distinct recipe has 22 videos. The procedure steps for each video are annotated with temporal boundaries and described by imperative English sentences (see the example below). The videos were downloaded from YouTube and are all in the third-person viewpoint. All the videos are unconstrained and can be performed by individual persons at their houses with unfixed cameras. YouCook2 contains rich recipe types and various cooking styles from all over the world.

Source: http://youcook2.eecs.umich.edu/ Image Source: https://competitions.codalab.org/competitions/20594

Benchmarks archive 2025-07-28

All 7 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

Papers archive 2025-07-28

30 shown of 33 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 198. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
HiCM²: Hierarchical Compact Memory Modeling for Dense Video Captioning 1 1 19 Dec 2024 not harvested
Do You Remember? Dense Video Captioning with Cross-Modal Memory Retrieval 1 1 11 Apr 2024 ran 2 of 3 samples (1 unverified)
MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding 1 1 8 Apr 2024 ran 7 of 9 samples (2 unverified)
Multi-granularity Correspondence Learning from Long-term Noisy Videos 1 2 30 Jan 2024 ran 2 of 3 samples (1 unverified)
OmniVec2 - A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning 0 1 1 Jan 2024 not harvested
OmniVec: Learning robust representations with cross modal sharing 0 2 7 Nov 2023 not harvested
HowToCaption: Prompting LLMs to Transform Video Annotations at Scale 1 3 7 Oct 2023 ran 4 of 5 samples (1 unverified; 5 pointer-only for licence)
COSA: Concatenated Sample Pretrained Vision-Language Foundation Model 1 1 15 Jun 2023 not harvested
VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset 2 2 29 May 2023 ran 15 of 42 samples (27 unverified)
MELTR: Meta Loss Transformer for Learning to Fine-tune Video Foundation Models 1 2 23 Mar 2023 ran 1 of 1 samples (0 unverified)
Text with Knowledge Graph Augmented Transformer for Video Captioning 0 1 22 Mar 2023 not harvested
Learning Grounded Vision-Language Representation for Versatile Understanding in Untrimmed Videos 1 1 11 Mar 2023 ran 1 of 12 samples (11 unverified)
Vid2Seq: Large-Scale Pretraining of a Visual Language Model for Dense Video Captioning 3 1 27 Feb 2023 not harvested
TempCLR: Temporal Alignment Representation with Contrastive Learning 1 1 28 Dec 2022 ran 3 of 5 samples (2 unverified)
VideoCoCa: Video-Text Modeling with Zero-Shot Transfer from Contrastive Captioners 0 3 9 Dec 2022 not harvested
OmniVL:One Foundation Model for Image-Language and Video-Language Tasks 0 1 15 Sep 2022 not harvested
Semantic Role Aware Correlation Transformer for Text to Video Retrieval 1 1 26 Jun 2022 not harvested
RoME: Role-aware Mixture-of-Expert Transformer for Text-to-Video Retrieval 1 1 26 Jun 2022 not harvested
MDMMT-2: Multidomain Multimodal Transformer for Video Retrieval, One More Step Towards Generalization 0 1 14 Mar 2022 not harvested
VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding 2 4 28 Sep 2021 not harvested
TACo: Token-aware Cascade Contrastive Learning for Video-Text Alignment 0 2 23 Aug 2021 not harvested
End-to-End Dense Video Captioning with Parallel Decoding 2 1 17 Aug 2021 ran 3 of 7 samples (4 unverified)
VLM: Task-agnostic Video-Language Model Pre-training for Video Understanding 1 2 20 May 2021 not harvested
Multimodal Clustering Networks for Self-supervised Learning from Unlabeled Videos 1 1 26 Apr 2021 ran 1 of 1 samples (0 unverified; 1 pointer-only for licence)
VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and Text 5 1 22 Apr 2021 ran 5 of 8 samples (3 unverified; 8 pointer-only for licence)
Multimodal Pretraining for Dense Video Captioning 1 2 10 Nov 2020 not harvested
COOT: Cooperative Hierarchical Transformer for Video-Text Representation Learning 1 2 1 Nov 2020 ran 0 of 6 samples (6 unverified)
UniVL: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation 2 2 15 Feb 2020 ran 1 of 4 samples (3 unverified; 1 pointer-only for licence)
End-to-End Learning of Visual Representations from Uncurated Instructional Videos 4 2 13 Dec 2019 ran 3 of 5 samples (2 unverified)
HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips 4 2 7 Jun 2019 ran 1 of 1 samples (0 unverified; 1 pointer-only for licence)

The full list of 33 is in the JSON twin.

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

Custom

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • YouCook2

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections