{"url":"/dataset/shot2story20k","name":"Shot2Story20K","full_name":null,"description_markdown":"A short clip of video may contain progression of multiple events and an interesting story line. A human needs to capture both the event in every shot and associate them together to understand the story behind it.\r\n\r\nIn this work, we present a new multi-shot video understanding benchmark Shot2Story with detailed shot-level captions and comprehensive video summaries. To facilitate better semantic understanding of videos, we provide captions for both visual signals and human narrations. We design several distinct tasks including single-shot video and narration captioning, multi-shot video summarization, and video retrieval with shot descriptions.\r\n\r\nPreliminary experiments show some challenges to generate a long and comprehensive video summary. Nevertheless, the generated imperfect summaries can already significantly boost the performance of existing video understanding tasks such as video question-answering, promoting an underexplored setting of video understanding with detailed summaries.","description_withheld":null,"homepage":"https://mingfei.info/shot2story/","introduced_date":"2023-12-16","introduced_date_note":null,"introduced_by":{"paper":"/paper/shot2story20k-a-new-benchmark-for","title":"Shot2Story20K: A New Benchmark for Comprehensive Understanding of Multi-shot Videos","first_author":"Mingfei Han","url":null},"license":null,"modalities":[{"name":"Videos","url":"/datasets/modality/videos"},{"name":"Texts","url":"/datasets/modality/texts"},{"name":"Audio","url":"/datasets/modality/audio"}],"tasks":[{"name":"Video Retrieval","url":"/task/video-retrieval","datasets_with_task":"/datasets/task/video-retrieval"},{"name":"Video Captioning","url":"/task/video-captioning","datasets_with_task":"/datasets/task/video-captioning"},{"name":"Zero-Shot Video Question Answer","url":"/task/zeroshot-video-question-answer","datasets_with_task":"/datasets/task/zeroshot-video-question-answer"},{"name":"Video Summarization","url":"/task/video-summarization","datasets_with_task":"/datasets/task/video-summarization"},{"name":"video narration captioning","url":"/task/video-narration-captioning","datasets_with_task":"/datasets/task/video-narration-captioning"}],"languages":[],"variants":["Shot2Story20K"],"data_loaders":[],"num_papers_in_archive":3,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/video-captioning-on-shot2story20k","task":"Video Captioning","dataset_variant":"Shot2Story20K","rows":2,"metrics":["CIDEr","BLEU-4","METEOR","ROUGE"],"first_row_in_archive_order":{"model":"Shotluck-Holmes (3.1B)","paper":"/paper/shotluck-holmes-a-family-of-efficient-small","metrics":{"BLEU-4":"8.7","CIDEr":"63.2","METEOR":"25.7","ROUGE":"36.2"},"code_links":[{"title":"Skyline-9/Shotluck-Holmes","url":"https://github.com/Skyline-9/Shotluck-Holmes"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/video-summarization-on-shot2story20k","task":"Video Summarization","dataset_variant":"Shot2Story20K","rows":2,"metrics":["CIDEr","BLEU-4","METEOR","ROUGE"],"first_row_in_archive_order":{"model":"Shotluck-Holmes (3.1B)","paper":"/paper/shotluck-holmes-a-family-of-efficient-small","metrics":{"BLEU-4":"7.67","CIDEr":"152.3","METEOR":"23.2","ROUGE":"43"},"code_links":[{"title":"Skyline-9/Shotluck-Holmes","url":"https://github.com/Skyline-9/Shotluck-Holmes"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/video-narration-captioning-on-shot2story20k","task":"video narration captioning","dataset_variant":"Shot2Story20K","rows":1,"metrics":["BLEU-4","CIDEr","METEOR","ROUGE"],"first_row_in_archive_order":{"model":"Ours","paper":"/paper/shot2story20k-a-new-benchmark-for","metrics":{"BLEU-4":"18.8","CIDEr":"168.7","METEOR":"24.8","ROUGE":"39"},"code_links":[{"title":"bytedance/Shot2Story","url":"https://github.com/bytedance/Shot2Story"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/shotluck-holmes-a-family-of-efficient-small","title":"Shotluck Holmes: A Family of Efficient Small-Scale Large Language Vision Models For Video Captioning and Summarization","date":"2024-05-31","rows_on_this_dataset":2,"code_links":1,"syntology":null},{"paper":"/paper/shot2story20k-a-new-benchmark-for","title":"Shot2Story20K: A New Benchmark for Comprehensive Understanding of Multi-shot Videos","date":"2023-12-16","rows_on_this_dataset":3,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":7,"samples_ran":6,"samples_unverified":1,"pointer_only_for_licence":7,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":1,"samples_harvested":7,"samples_ran":6,"samples_unverified":1,"pointer_only_for_licence":7,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}