{"url":"/dataset/youcook","name":"YouCook","full_name":null,"description_markdown":"This data set was prepared from 88 open-source YouTube cooking videos. The YouCook dataset contains videos of people cooking various recipes. The videos were downloaded from YouTube and are all in the third-person viewpoint; they represent a significantly more challenging visual problem than existing cooking and kitchen datasets (the background kitchen/scene is different for many and most videos have dynamic camera changes). In addition, frame-by-frame object and action annotations are provided for training data (as well as a number of precomputed low-level features). Finally, each video has a number of human provided natural language descriptions (on average, there are eight different descriptions per video). This dataset has been created to serve as a benchmark in describing complex real-world videos with natural language descriptions.\r\n\r\nSource: [A Thousand Frames in Just a Few Words: Lingual Description of Videos through Latent Topics and Sparse Object Stitching](/paper/a-thousand-frames-in-just-a-few-words-lingual)","description_withheld":null,"homepage":"https://web.eecs.umich.edu/~jjcorso/r/youcook/","introduced_date":null,"introduced_date_note":null,"introduced_by":{"paper":"/paper/a-thousand-frames-in-just-a-few-words-lingual","title":"A Thousand Frames in Just a Few Words: Lingual Description of Videos through Latent Topics and Sparse Object Stitching","first_author":"Pradipto Das","url":null},"license":{"name":"Unknown","url":null},"modalities":[{"name":"Videos","url":"/datasets/modality/videos"}],"tasks":[{"name":"Video Captioning","url":"/task/video-captioning","datasets_with_task":"/datasets/task/video-captioning"},{"name":"Dense Video Captioning","url":"/task/dense-video-captioning","datasets_with_task":"/datasets/task/dense-video-captioning"},{"name":"Video Description","url":"/task/video-description","datasets_with_task":"/datasets/task/video-description"}],"languages":[],"variants":["YouCook"],"data_loaders":[],"num_papers_in_archive":45,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}