{"url":"/dataset/vitt","name":"ViTT","full_name":"Video Timeline Tags","description_markdown":"The ViTT dataset consists of human produced segment-level annotations for 8,169 videos. Of these, 5,840 videos have been annotated once, and the rest of the videos have been annotated twice or more. A total of 12,461 sets of annotations are released. The videos in the dataset are from the [Youtube-8M dataset](https://paperswithcode.com/dataset/youtube-8m).\r\n\r\nAn annotation has the following format:\r\n```\r\n{\r\n  \"id\": \"FmTp\",\r\n  \"annotations\": [\r\n    {\r\n      \"timestamp\": 260,\r\n      \"tag\": \"Opening\"\r\n    },\r\n    {\r\n      \"timestamp\": 16000,\r\n      \"tag\": \"Displaying technique\"\r\n    },\r\n    {\r\n      \"timestamp\": 23990,\r\n      \"tag\": \"Showing foot positioning\"\r\n    },\r\n    {\r\n      \"timestamp\": 55530,\r\n      \"tag\": \"Demonstrating crossover\"\r\n    },\r\n    {\r\n      \"timestamp\": 114100,\r\n      \"tag\": \"Closing\"\r\n    }\r\n  ]\r\n}\r\n```\r\n\r\nSource: [Video Timeline Tags (ViTT)](https://github.com/google-research-datasets/Video-Timeline-Tags-ViTT)","description_withheld":null,"homepage":"https://github.com/google-research-datasets/Video-Timeline-Tags-ViTT","introduced_date":null,"introduced_date_note":null,"introduced_by":{"paper":"/paper/multimodal-pretraining-for-dense-video","title":"Multimodal Pretraining for Dense Video Captioning","first_author":"Gabriel Huang","url":null},"license":null,"modalities":[{"name":"Videos","url":"/datasets/modality/videos"}],"tasks":[{"name":"Video Captioning","url":"/task/video-captioning","datasets_with_task":"/datasets/task/video-captioning"},{"name":"Dense Video Captioning","url":"/task/dense-video-captioning","datasets_with_task":"/datasets/task/dense-video-captioning"},{"name":"Zero-shot dense video captioning","url":"/task/zero-shot-dense-video-captioning","datasets_with_task":"/datasets/task/zero-shot-dense-video-captioning"}],"languages":[],"variants":["ViTT"],"data_loaders":[{"repo":"https://github.com/google-research-datasets/Video-Timeline-Tags-ViTT","url":"https://github.com/google-research-datasets/Video-Timeline-Tags-ViTT","frameworks":[]}],"num_papers_in_archive":14,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/dense-video-captioning-on-vitt","task":"Dense Video Captioning","dataset_variant":"ViTT","rows":4,"metrics":["SODA","CIDEr","METEOR"],"first_row_in_archive_order":{"model":"Vid2Seq (VidChapters-7M PT)","paper":null,"metrics":{"CIDEr":"50.9","METEOR":"9.5","SODA":"0.151"},"code_links":[]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/zero-shot-dense-video-captioning-on-vitt","task":"Zero-shot dense video captioning","dataset_variant":"ViTT","rows":1,"metrics":["CIDEr","METEOR","SODA"],"first_row_in_archive_order":{"model":"Vid2Seq (VidChapters-7M PT)","paper":null,"metrics":{"CIDEr":"30.2","METEOR":"6.7","SODA":"9.1"},"code_links":[]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/hicm-2-hierarchical-compact-memory-modeling","title":"HiCM$^2$: Hierarchical Compact Memory Modeling for Dense Video Captioning","date":"2024-12-19","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/vid2seq-large-scale-pretraining-of-a-visual","title":"Vid2Seq: Large-Scale Pretraining of a Visual Language Model for Dense Video Captioning","date":"2023-02-27","rows_on_this_dataset":1,"code_links":3,"syntology":null},{"paper":"/paper/end-to-end-dense-video-captioning-as-sequence-1","title":"End-to-end Dense Video Captioning as Sequence Generation","date":"2022-04-18","rows_on_this_dataset":1,"code_links":0,"syntology":null}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}