{"url":"/task/video-description","name":"Video Description","slug":"video-description","description_markdown":"The goal of automatic **Video Description** is to tell a story about events happening in a video. While early Video Description methods produced captions for short clips that were manually segmented to contain a single event of interest, more recently dense video captioning has been proposed to both segment distinct events in time and describe them in a series of coherent sentences. This problem is a generalization of dense image region captioning and has many practical applications, such as generating textual summaries for the visually impaired, or detecting and describing important events in surveillance footage.\r\n\r\n\r\n<span class=\"description-source\">Source: [Joint Event Detection and Description in Continuous Video Streams ](https://arxiv.org/abs/1802.10250)</span>","categories":[{"name":"Computer Vision","url":"/area/computer-vision"}],"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","slug_source":"archive_url"},"counts":{"papers_tagged":104,"papers_with_code":34,"benchmarks":0,"benchmark_tables_in_archive":0,"benchmark_tables_shown":0,"benchmark_tables_withheld_as_spam":0,"benchmark_definition":"a leaderboard table with at least one row; benchmark_tables_shown also counts the zero-row tables; benchmark_tables_in_archive adds the tables withheld as spam","datasets":9,"subtasks":0,"parent_tasks":1},"benchmarks":[],"datasets":[{"url":"/dataset/flickr30k","name":"Flickr30k","full_name":"Flickr30k","num_papers_in_archive":880},{"url":"/dataset/tacos-multi-level-corpus","name":"TACoS Multi-Level Corpus","full_name":"","num_papers_in_archive":45},{"url":"/dataset/youcook","name":"YouCook","full_name":"","num_papers_in_archive":45},{"url":"/dataset/activitynet-entities-1","name":"ActivityNet Entities","full_name":null,"num_papers_in_archive":18},{"url":"/dataset/videocc3m","name":"VideoCC3M","full_name":"Video-Conceptual-Captions","num_papers_in_archive":10},{"url":"/dataset/edub-seg","name":"EDUB-Seg","full_name":"Egocentric Dataset of the University of Barcelona – Segmentation","num_papers_in_archive":4},{"url":"/dataset/m-vad-names","name":"M-VAD Names","full_name":"M-VAD Names Dataset","num_papers_in_archive":3},{"url":"/dataset/devan","name":"DeVAn","full_name":"Dense Video Annotation for Video-Language Models","num_papers_in_archive":1},{"url":"/dataset/vidas","name":"ViDAS","full_name":"","num_papers_in_archive":1}],"subtasks":[],"parent_tasks":[{"url":"/task/video","name":"Video"}],"papers":{"order":"repositories listed in the archive (desc), then date (desc); the archive holds no stars","population":"papers tagged with this task that list at least one repository in the archive","shown":30,"of":34,"tagged_in_all":104,"items":[{"url":"/paper/describing-videos-by-exploiting-temporal","title":"Describing Videos by Exploiting Temporal Structure","date":"2015-02-27","arxiv_id":"1502.08029","repositories_listed":5,"syntology":{"n":2,"n_ran":0,"n_unverified":2,"n_pointer_only":0}},{"url":"/paper/vatex-a-large-scale-high-quality-multilingual","title":"VATEX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language Research","date":"2019-04-06","arxiv_id":"1904.03493","repositories_listed":4,"syntology":null},{"url":"/paper/audio-visual-scene-aware-dialog-avsd","title":"Audio Visual Scene-Aware Dialog (AVSD) Challenge at DSTC7","date":"2018-06-01","arxiv_id":"1806.00525","repositories_listed":4,"syntology":null},{"url":"/paper/improving-lstm-based-video-description-with","title":"Improving LSTM-based Video Description with Linguistic Knowledge Mined from Text","date":"2016-04-06","arxiv_id":"1604.01729","repositories_listed":3,"syntology":null},{"url":"/paper/grounded-video-description","title":"Grounded Video Description","date":"2018-12-17","arxiv_id":"1812.06587","repositories_listed":2,"syntology":null},{"url":"/paper/end-to-end-audio-visual-scene-aware-dialog","title":"End-to-End Audio Visual Scene-Aware Dialog using Multimodal Attention-Based Video Features","date":"2018-06-21","arxiv_id":"1806.08409","repositories_listed":2,"syntology":{"n":3,"n_ran":0,"n_unverified":3,"n_pointer_only":0}},{"url":"/paper/tarsier2-advancing-large-vision-language","title":"Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding","date":"2025-01-14","arxiv_id":"2501.07888","repositories_listed":1,"syntology":{"n":4,"n_ran":4,"n_unverified":0,"n_pointer_only":4}},{"url":"/paper/implicit-location-caption-alignment-via","title":"Implicit Location-Caption Alignment via Complementary Masking for Weakly-Supervised Dense Video Captioning","date":"2024-12-17","arxiv_id":"2412.12791","repositories_listed":1,"syntology":{"n":12,"n_ran":0,"n_unverified":12,"n_pointer_only":0}},{"url":"/paper/storyteller-improving-long-video-description","title":"StoryTeller: Improving Long Video Description through Global Audio-Visual Character Identification","date":"2024-11-11","arxiv_id":"2411.07076","repositories_listed":1,"syntology":null},{"url":"/paper/sustechgan-image-generation-for-object","title":"SUSTechGAN: Image Generation for Object Detection in Adverse Conditions of Autonomous Driving","date":"2024-07-18","arxiv_id":"2408.01430","repositories_listed":1,"syntology":null},{"url":"/paper/tarsier-recipes-for-training-and-evaluating-1","title":"Tarsier: Recipes for Training and Evaluating Large Video Description Models","date":"2024-06-30","arxiv_id":"2407.00634","repositories_listed":1,"syntology":{"n":2,"n_ran":2,"n_unverified":0,"n_pointer_only":0}},{"url":"/paper/hawk-learning-to-understand-open-world-video","title":"Hawk: Learning to Understand Open-World Video Anomalies","date":"2024-05-27","arxiv_id":"2405.16886","repositories_listed":1,"syntology":{"n":9,"n_ran":7,"n_unverified":2,"n_pointer_only":9}},{"url":"/paper/trafficvlm-a-controllable-visual-language","title":"TrafficVLM: A Controllable Visual Language Model for Traffic Video Captioning","date":"2024-04-14","arxiv_id":"2404.09275","repositories_listed":1,"syntology":{"n":1,"n_ran":1,"n_unverified":0,"n_pointer_only":0}},{"url":"/paper/jmi-at-semeval-2024-task-3-two-step-approach","title":"JMI at SemEval 2024 Task 3: Two-step approach for multimodal ECAC using in-context learning with GPT and instruction-tuned Llama models","date":"2024-03-05","arxiv_id":"2403.04798","repositories_listed":1,"syntology":{"n":12,"n_ran":4,"n_unverified":8,"n_pointer_only":12}},{"url":"/paper/panda-70m-captioning-70m-videos-with-multiple","title":"Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers","date":"2024-02-29","arxiv_id":"2402.19479","repositories_listed":1,"syntology":null},{"url":"/paper/funqa-towards-surprising-video-comprehension","title":"FunQA: Towards Surprising Video Comprehension","date":"2023-06-26","arxiv_id":"2306.14899","repositories_listed":1,"syntology":{"n":3,"n_ran":0,"n_unverified":3,"n_pointer_only":0}},{"url":"/paper/msvd-indonesian-a-benchmark-for-multimodal","title":"MSVD-Indonesian: A Benchmark for Multimodal Video-Text Tasks in Indonesian","date":"2023-06-20","arxiv_id":"2306.11341","repositories_listed":1,"syntology":null},{"url":"/paper/edit-as-you-wish-video-description-editing","title":"Edit As You Wish: Video Caption Editing with Multi-grained User Control","date":"2023-05-15","arxiv_id":"2305.08389","repositories_listed":1,"syntology":null},{"url":"/paper/fine-grained-audible-video-description","title":"Fine-grained Audible Video Description","date":"2023-03-27","arxiv_id":"2303.15616","repositories_listed":1,"syntology":{"n":12,"n_ran":0,"n_unverified":12,"n_pointer_only":0}},{"url":"/paper/thinking-hallucination-for-video-captioning","title":"Thinking Hallucination for Video Captioning","date":"2022-09-28","arxiv_id":"2209.13853","repositories_listed":1,"syntology":null},{"url":"/paper/what-s-in-a-caption-dataset-specific","title":"What's in a Caption? Dataset-Specific Linguistic Diversity and Its Effect on Visual Description Models and Metrics","date":"2022-05-12","arxiv_id":"2205.06253","repositories_listed":1,"syntology":null},{"url":"/paper/learn-to-understand-negation-in-video","title":"Learn to Understand Negation in Video Retrieval","date":"2022-04-30","arxiv_id":"2205.00132","repositories_listed":1,"syntology":null},{"url":"/paper/identity-aware-multi-sentence-video","title":"Identity-Aware Multi-Sentence Video Description","date":"2020-08-22","arxiv_id":"2008.09791","repositories_listed":1,"syntology":null},{"url":"/paper/describing-unseen-videos-via-multi","title":"Describing Unseen Videos via Multi-Modal Cooperative Dialog Agents","date":"2020-08-18","arxiv_id":"2008.07935","repositories_listed":1,"syntology":null},{"url":"/paper/delving-deeper-into-the-decoder-for-video","title":"Delving Deeper into the Decoder for Video Captioning","date":"2020-01-16","arxiv_id":"2001.05614","repositories_listed":1,"syntology":{"n":2,"n_ran":0,"n_unverified":2,"n_pointer_only":0}},{"url":"/paper/vizseq-a-visual-analysis-toolkit-for-text","title":"VizSeq: A Visual Analysis Toolkit for Text Generation Tasks","date":"2019-09-12","arxiv_id":"1909.05424","repositories_listed":1,"syntology":null},{"url":"/paper/adversarial-inference-for-multi-sentence","title":"Adversarial Inference for Multi-Sentence Video Description","date":"2018-12-13","arxiv_id":"1812.05634","repositories_listed":1,"syntology":null},{"url":"/paper/predicting-visual-features-from-text-for","title":"Predicting Visual Features from Text for Image and Video Caption Retrieval","date":"2017-09-05","arxiv_id":"1709.01362","repositories_listed":1,"syntology":null},{"url":"/paper/egocentric-video-description-based-on","title":"Egocentric Video Description based on Temporally-Linked Sequences","date":"2017-04-07","arxiv_id":"1704.02163","repositories_listed":1,"syntology":null},{"url":"/paper/memory-augmented-attention-modelling-for","title":"Memory-augmented Attention Modelling for Videos","date":"2016-11-07","arxiv_id":"1611.02261","repositories_listed":1,"syntology":null}],"syntology_records":11,"syntology_note":"a paper without a record is not a recorded non-run: it may lack an arXiv id or simply be absent from the graph layer"},"description_links":{"kept":0,"unwrapped_to_text":0,"bare_urls_linked":0,"relative_images_dropped":0,"rule":"internal links are kept only when the target slug exists in the catalog"},"syntology":{"read_at":"2026-09-24T18:15:14+00:00","claim":"Per-sample execution status on synthesized fixtures ('ran N of M samples'); not a correctness claim and not a ranking signal.","status_vocabulary":{"ran_honours":"ran, honoured the contract we drafted","ran_violates":"ran, violated the contract we drafted","ran_draft_wrong":"ran; our contract draft was wrong, not the code","ran_fixture":"ran; our fixture could not drive it","ran":"ran on a synthesized input","unverified":"unverified (harvested, no recorded run)"}},"not_shown":{"libraries":"the archive has no per-task library table","trend_sparklines":"the Trend column of the benchmarks table was a rendered image; it is not in the archive","social_and_latest_sorts":"stars and social signals are not in the archive"}}