Datasets › LSMDC

LSMDC (Large Scale Movie Description Challenge)

Introduced by Anna Rohrbach et al. in A Dataset for Movie Description1 Jan 2015 archive 2025-07-28

This dataset contains 118,081 short video clips extracted from 202 movies. Each video has a caption, either extracted from the movie script or from transcribed DVS (descriptive video services) for the visually impaired. The validation set contains 7408 clips and evaluation is performed on a test set of 1000 videos from movies disjoint from the training and val sets.

Source: Use What You Have: Video Retrieval Using Representations From Collaborative Experts Image Source: https://sites.google.com/site/describingmovies/

Benchmarks archive 2025-07-28

All 3 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Video Retrieval LSMDC InternVideo2-6B text-to-video R@1 46.4 InternVideo2: Scaling Foundation Models for Multimodal... opengvlab/internvideo +1 38 Compare
Zero-Shot Video Retrieval LSMDC InternVideo2-6B text-to-video R@1 33.8 InternVideo2: Scaling Foundation Models for Multimodal... opengvlab/internvideo +1 16 Compare
Zero-Shot Learning LSMDC FrozenBiLM Accuracy 51.5 Zero-Shot Video Question Answering via Frozen... antoyang/FrozenBiLM +2 1 Compare

Papers archive 2025-07-28

30 shown of 41 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 126. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
InternVideo2: Scaling Foundation Models for Multimodal Video Understanding 2 3 22 Mar 2024 not harvested
vid-TLDR: Training Free Token merging for Light-weight Video Transformer 1 1 20 Mar 2024 ran 3 of 4 samples (1 unverified)
HowToCaption: Prompting LLMs to Transform Video Annotations at Scale 1 2 7 Oct 2023 ran 4 of 5 samples (1 unverified; 5 pointer-only for licence)
BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning 1 1 27 Sep 2023 ran 1 of 4 samples (3 unverified)
COSA: Concatenated Sample Pretrained Vision-Language Foundation Model 1 1 15 Jun 2023 not harvested
VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and Dataset 1 1 17 Apr 2023 not harvested
Unmasked Teacher: Towards Training-Efficient Video Foundation Models 1 2 28 Mar 2023 ran 3 of 8 samples (5 unverified)
DiffusionRet: Generative Text-Video Retrieval with Diffusion Model 4 1 17 Mar 2023 ran 5 of 6 samples (1 unverified)
mPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and Video 4 2 1 Feb 2023 ran 9 of 19 samples (10 unverified)
Revisiting Temporal Modeling for CLIP-based Image-to-Video Knowledge Transferring 1 1 26 Jan 2023 not harvested
HiTeA: Hierarchical Temporal-Aware Video-Language Pre-training 0 3 30 Dec 2022 not harvested
InternVideo: General Video Foundation Models via Generative and Discriminative Learning 2 2 6 Dec 2022 ran 3 of 3 samples (0 unverified)
Seeing What You Miss: Vision-Language Pre-training with Semantic Completion Learning 1 1 24 Nov 2022 not harvested
Expectation-Maximization Contrastive Learning for Compact Video-and-Language Representations 4 3 21 Nov 2022 ran 3 of 4 samples (1 unverified)
CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Representation Alignment 1 1 14 Sep 2022 ran 4 of 4 samples (0 unverified; 4 pointer-only for licence)
An Empirical Study of End-to-End Video-Language Transformers with Masked Visual Modeling 1 1 4 Sep 2022 not harvested
Clover: Towards A Unified Video-Language Alignment and Fusion Model 1 2 16 Jul 2022 not harvested
X-CLIP: End-to-End Multi-grained Contrastive Learning for Video-Text Retrieval 3 1 15 Jul 2022 ran 1 of 2 samples (1 unverified; 1 pointer-only for licence)
Zero-Shot Video Question Answering via Frozen Bidirectional Language Models 3 1 16 Jun 2022 ran 14 of 34 samples (20 unverified; 1 pointer-only for licence)
CenterCLIP: Token Clustering for Efficient Text-Video Retrieval 1 1 2 May 2022 ran 0 of 1 samples (1 unverified; 1 pointer-only for licence)
MILES: Visual BERT Pre-training with Injected Language Semantics for Video-text Retrieval 1 1 26 Apr 2022 not harvested
Tencent Text-Video Retrieval: Hierarchical Cross-Modal Interactions with Multi-Level Representations 0 2 7 Apr 2022 not harvested
X-Pool: Cross-Modal Language-Video Attention for Text-Video Retrieval 1 1 28 Mar 2022 ran 1 of 2 samples (1 unverified; 2 pointer-only for licence)
MDMMT-2: Multidomain Multimodal Transformer for Video Retrieval, One More Step Towards Generalization 0 1 14 Mar 2022 not harvested
Bridging Video-text Retrieval with Multiple Choice Questions 2 1 13 Jan 2022 ran 13 of 24 samples (11 unverified; 6 pointer-only for licence)
Cross Modal Retrieval with Querybank Normalisation 1 1 23 Dec 2021 ran 2 of 3 samples (1 unverified)
Advancing High-Resolution Video-Language Representation with Large-Scale Video Transcriptions 1 1 19 Nov 2021 not harvested
Video and Text Matching with Conditioned Embeddings 1 1 21 Oct 2021 not harvested
Improving Video-Text Retrieval by Multi-Stream Corpus Alignment and Dual Softmax Loss 2 1 9 Sep 2021 not harvested
CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval 5 2 18 Apr 2021 ran 3 of 4 samples (1 unverified; 3 pointer-only for licence)

The full list of 41 is in the JSON twin.

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

Custom

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • LSMDC

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections