{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/dense-captioning-events-in-videos","title":"Dense-Captioning Events in Videos","arxiv_id":"1705.00754","date":"2017-05-02","proceeding":"ICCV 2017 10","authors":["Ranjay Krishna","Kenji Hata","Frederic Ren","Li Fei-Fei","Juan Carlos Niebles"],"abstract":"Most natural videos contain numerous events. For example, in a video of a\n\"man playing a piano\", the video might also contain \"another man dancing\" or \"a\ncrowd clapping\". We introduce the task of dense-captioning events, which\ninvolves both detecting and describing events in a video. We propose a new\nmodel that is able to identify all events in a single pass of the video while\nsimultaneously describing the detected events with natural language. Our model\nintroduces a variant of an existing proposal module that is designed to capture\nboth short as well as long events that span minutes. To capture the\ndependencies between the events in a video, our model introduces a new\ncaptioning module that uses contextual information from past and future events\nto jointly describe all events. We also introduce ActivityNet Captions, a\nlarge-scale benchmark for dense-captioning events. ActivityNet Captions\ncontains 20k videos amounting to 849 video hours with 100k total descriptions,\neach with it's unique start and end time. Finally, we report performances of\nour model for dense-captioning events, video retrieval and localization.","url_abs":"http://arxiv.org/abs/1705.00754v1","url_pdf":"http://arxiv.org/pdf/1705.00754v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"dense-captioning-events-in-videos","repo_url":"https://github.com/HanNight/RE-T5","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"Apache-2.0"}},{"paper_slug":"dense-captioning-events-in-videos","repo_url":"https://github.com/akoepke/audio-retrieval-benchmark","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"dense-captioning-events-in-videos","repo_url":"https://github.com/oncescuandreea/audio-retrieval","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"Apache-2.0"}},{"paper_slug":"dense-captioning-events-in-videos","repo_url":"https://github.com/sangminwoo/explore-and-match","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"dense-captioning","task_name":"Dense Captioning"},{"task_slug":"retrieval","task_name":"Retrieval"},{"task_slug":"video-retrieval","task_name":"Video Retrieval"}],"methods":[],"datasets_introduced":[{"slug":"activitynet-captions","name":"ActivityNet Captions","full_name":""}],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1705.00754","atlas_url":"https://app.syntology.ai/?focus=1705.00754","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}