Papers › Improving LSTM-based Video Description with Linguistic Knowledge Mined from Text

Improving LSTM-based Video Description with Linguistic Knowledge Mined from Text

6 Apr 2016EMNLP 2016 11arXiv:1604.01729archive 2025-07-28

Subhashini Venugopalan, Lisa Anne Hendricks, Raymond Mooney, Kate Saenko

This paper investigates how linguistic knowledge mined from large text corpora can aid the generation of natural language descriptions of videos. Specifically, we integrate both a neural language model and distributional semantics trained on large text corpora into a recent LSTM-based architecture for video description. We evaluate our approach on a collection of Youtube videos as well as two large movie description datasets showing significant improvements in grammaticality while modestly improving descriptive quality.

PaperPDFConference PDFCode

Code

TejInaco/multimodalML mentioned on GitHub report
beerzyp/ECAC-Chain-Fusion mentioned on GitHubtf report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

DescriptiveLanguage ModelingLanguage ModellingVideo Description

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections