Browse State-of-the-Art › Video Description
Video Description
34 papers with code · 0 benchmarks · 9 datasets archive 2025-07-28
The goal of automatic Video Description is to tell a story about events happening in a video. While early Video Description methods produced captions for short clips that were manually segmented to contain a single event of interest, more recently dense video captioning has been proposed to both segment distinct events in time and describe them in a series of coherent sentences. This problem is a generalization of dense image region captioning and has many practical applications, such as generating textual summaries for the visually impaired, or detecting and describing important events in surveillance footage.
Source: Joint Event Detection and Description in Continuous Video Streams
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
9 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 34 papers with code (104 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
27 Feb 2015 5 repositories listed Syntology ran 0 of 2 samples · 2 unverifiedIn this context, we propose an approach that successfully takes into account both the local and global temporal structure of videos to produce descriptions.
-
6 Apr 2019 4 repositories listedWe also introduce two tasks for video-and-language research based on VATEX: (1) Multilingual Video Captioning, aimed at describing a video in various languages with a compact unified captioning model, and (2)…
-
1 Jun 2018 4 repositories listedScene-aware dialog systems will be able to have conversations with users about the objects and events around them.
-
6 Apr 2016 3 repositories listedThis paper investigates how linguistic knowledge mined from large text corpora can aid the generation of natural language descriptions of videos.
-
17 Dec 2018 2 repositories listedOur dataset, ActivityNet-Entities, augments the challenging ActivityNet Captions dataset with 158k bounding box annotations, each grounding a noun phrase.
-
21 Jun 2018 2 repositories listed Syntology ran 0 of 3 samples · 3 unverifiedWe introduce a new dataset of dialogs about videos of human behaviors.
-
14 Jan 2025 1 repository listed Syntology ran 4 of 4 samples · 0 unverified · 4 pointer-only (licence)We introduce Tarsier2, a state-of-the-art large vision-language model (LVLM) designed for generating detailed and accurate video descriptions, while also exhibiting superior general video understanding capabilities.
-
17 Dec 2024 1 repository listed Syntology ran 0 of 12 samples · 12 unverifiedExisting methods rely on explicit alignment constraints between event locations and captions, which involve complex event proposal procedures during both training and inference.
-
11 Nov 2024 1 repository listedWe propose StoryTeller, a system for generating dense descriptions of long videos, incorporating both low-level visual concepts and high-level plot information.
-
18 Jul 2024 1 repository listedWe test the SUSTechGAN and the well-known GANs to generate driving images in adverse conditions of rain and night and apply the generated images to retrain object detection networks.
-
30 Jun 2024 1 repository listed Syntology ran 2 of 2 samples · 0 unverifiedOur second contribution is the introduction of a new benchmark -- DREAM-1K (https://tarsier-vlm.
-
27 May 2024 1 repository listed Syntology ran 7 of 9 samples · 2 unverified · 9 pointer-only (licence)Video Anomaly Detection (VAD) systems can autonomously monitor and identify disturbances, reducing the need for manual labor and associated costs.
-
14 Apr 2024 1 repository listed Syntology ran 1 of 1 samples · 0 unverifiedTraffic video description and analysis have received much attention recently due to the growing demand for efficient and reliable urban surveillance systems.
-
5 Mar 2024 1 repository listed Syntology ran 4 of 12 samples · 8 unverified · 12 pointer-only (licence)However, the complexities of these diverse modalities pose challenges for developing an efficient multimodal emotion cause analysis (ECA) system.
-
29 Feb 2024 1 repository listedNext, we finetune a retrieval model on a small subset where the best caption of each video is manually selected and then employ the model in the whole dataset to select the best caption as the annotation.
-
26 Jun 2023 1 repository listed Syntology ran 0 of 3 samples · 3 unverifiedSurprising videos, such as funny clips, creative performances, or visual illusions, attract significant attention.
-
20 Jun 2023 1 repository listedSince the availability of the pretraining resources with Indonesian sentences is relatively limited, the applicability of those approaches to our dataset is still questionable.
-
15 May 2023 1 repository listedIn this paper, we propose a novel \textbf{V}ideo \textbf{C}aption \textbf{E}diting \textbf{(VCE)} task to automatically revise an existing video description guided by multi-grained user requests.
-
27 Mar 2023 1 repository listed Syntology ran 0 of 12 samples · 12 unverifiedWe explore a new task for audio-visual-language modeling called fine-grained audible video description (FAVD).
-
28 Sep 2022 1 repository listedIn video captioning, there are two kinds of hallucination: object and action hallucination.
-
12 May 2022 1 repository listedWhile there have been significant gains in the field of automated video description, the generalization performance of automated description models to novel domains remains a major barrier to using these systems in the…
-
30 Apr 2022 1 repository listedWe propose a learning based method for training a negation-aware video retrieval model.
-
22 Aug 2020 1 repository listedThis auxiliary task allows us to propose a two-stage approach to Identity-Aware Video Description.
-
18 Aug 2020 1 repository listedWith the arising concerns for the AI systems provided with direct access to abundant sensitive information, researchers seek to develop more reliable AI with implicit information sources.
-
16 Jan 2020 1 repository listed Syntology ran 0 of 2 samples · 2 unverifiedVideo captioning is an advanced multi-modal task which aims to describe a video clip using a natural language sentence.
-
12 Sep 2019 1 repository listedAutomatic evaluation of text generation tasks (e.
-
13 Dec 2018 1 repository listedAmong the main issues are the fluency and coherence of the generated descriptions, and their relevance to the video.
-
5 Sep 2017 1 repository listedThis paper strives to find amidst a set of sentences the one best describing the content of a given image or video.
-
7 Apr 2017 1 repository listedWe propose a novel methodology that exploits information from temporally neighboring events, matching precisely the nature of egocentric sequences.
-
7 Nov 2016 1 repository listedWe present a method to improve video description generation by modeling higher-order interactions between video frames and described concepts.
Syntology lines on 11 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections