Datasets › Shot2Story20K
Shot2Story20K
A short clip of video may contain progression of multiple events and an interesting story line. A human needs to capture both the event in every shot and associate them together to understand the story behind it.
In this work, we present a new multi-shot video understanding benchmark Shot2Story with detailed shot-level captions and comprehensive video summaries. To facilitate better semantic understanding of videos, we provide captions for both visual signals and human narrations. We design several distinct tasks including single-shot video and narration captioning, multi-shot video summarization, and video retrieval with shot descriptions.
Preliminary experiments show some challenges to generate a long and comprehensive video summary. Nevertheless, the generated imperfect summaries can already significantly boost the performance of existing video understanding tasks such as video question-answering, promoting an underexplored setting of video understanding with detailed summaries.
Benchmarks archive 2025-07-28
All 3 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.
| First row (archive order) | Paper | Code | ||||
|---|---|---|---|---|---|---|
| Video Captioning | Shot2Story20K | Shotluck-Holmes (3.1B) CIDEr 63.2 | Shotluck Holmes: A Family of Efficient Small-Scale Large... | Skyline-9/Shotluck-Holmes | 2 | Compare |
| Video Summarization | Shot2Story20K | Shotluck-Holmes (3.1B) CIDEr 152.3 | Shotluck Holmes: A Family of Efficient Small-Scale Large... | Skyline-9/Shotluck-Holmes | 2 | Compare |
| video narration captioning | Shot2Story20K | Ours BLEU-4 18.8 | Shot2Story20K: A New Benchmark for Comprehensive... | bytedance/Shot2Story | 1 | Compare |
Papers archive 2025-07-28
2 shown of 2 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 3. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.
| Date | Samples run Syntology | |||
|---|---|---|---|---|
| Shotluck Holmes: A Family of Efficient Small-Scale Large Language Vision Models For Video Captioning and Summarization | 1 | 2 | 31 May 2024 | not harvested |
| Shot2Story20K: A New Benchmark for Comprehensive Understanding of Multi-shot Videos | 1 | 3 | 16 Dec 2023 | ran 6 of 7 samples (1 unverified; 7 pointer-only for licence) |
Dataset loaders archive 2025-07-28
No loader listed in the archive.
Tasks archive 2025-07-28
License archive 2025-07-28
No licence recorded in the archive. Absence here is not a statement about the dataset's terms.
Modalities archive 2025-07-28
Languages archive 2025-07-28
No language tagged.
Variants archive 2025-07-28
- Shot2Story20K
1 variant name, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections