{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/streamlined-dense-video-captioning","title":"Streamlined Dense Video Captioning","arxiv_id":"1904.03870","date":"2019-04-08","proceeding":"CVPR 2019 6","authors":["Jonghwan Mun","Linjie Yang","Zhou Ren","Ning Xu","Bohyung Han"],"abstract":"Dense video captioning is an extremely challenging task since accurate and\ncoherent description of events in a video requires holistic understanding of\nvideo contents as well as contextual reasoning of individual events. Most\nexisting approaches handle this problem by first detecting event proposals from\na video and then captioning on a subset of the proposals. As a result, the\ngenerated sentences are prone to be redundant or inconsistent since they fail\nto consider temporal dependency between events. To tackle this challenge, we\npropose a novel dense video captioning framework, which models temporal\ndependency across events in a video explicitly and leverages visual and\nlinguistic context from prior events for coherent storytelling. This objective\nis achieved by 1) integrating an event sequence generation network to select a\nsequence of event proposals adaptively, and 2) feeding the sequence of event\nproposals to our sequential video captioning network, which is trained by\nreinforcement learning with two-level rewards at both event and episode levels\nfor better context modeling. The proposed technique achieves outstanding\nperformances on ActivityNet Captions dataset in most metrics.","url_abs":"http://arxiv.org/abs/1904.03870v1","url_pdf":"http://arxiv.org/pdf/1904.03870v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"streamlined-dense-video-captioning","repo_url":"https://github.com/ttengwang/ESGN","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"dense-video-captioning","task_name":"Dense Video Captioning"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"video-captioning","task_name":"Video Captioning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1904.03870","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}