Datasets › Video Localized Narratives

Video Localized Narratives

Introduced by Paul Voigtlaender et al. in Connecting Vision and Language with Video Localized Narratives22 Feb 2023 archive 2025-07-28

Video Localized Narratives is a new form of multimodal video annotations connecting vision and language. The annotations are created from videos with Localized Narratives, capturing even complex events involving multiple actors interacting with each other and with several passive objects. It contains annotations of 20k videos of the OVIS, UVO, and Oops datasets, totalling 1.7M words.

Source: Connecting Vision and Language with Video Localized Narratives

Image Source: https://arxiv.org/pdf/2302.11217v1.pdf

Benchmarks archive 2025-07-28

No leaderboard in the archive resolves to this dataset.

Papers archive 2025-07-28

No paper in the archive has a leaderboard row on this dataset; the archive counts 3 papers for it but never published that list.

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

CC BY 4.0 License

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • Video Localized Narratives

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections