{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/coherent-multi-sentence-video-description","title":"Coherent Multi-Sentence Video Description with Variable Level of Detail","arxiv_id":"1403.6173","date":"2014-03-24","proceeding":null,"authors":["Anna Senina","Marcus Rohrbach","Wei Qiu","Annemarie Friedrich","Sikandar Amin","Mykhaylo Andriluka","Manfred Pinkal","Bernt Schiele"],"abstract":"Humans can easily describe what they see in a coherent way and at varying\nlevel of detail. However, existing approaches for automatic video description\nare mainly focused on single sentence generation and produce descriptions at a\nfixed level of detail. In this paper, we address both of these limitations: for\na variable level of detail we produce coherent multi-sentence descriptions of\ncomplex videos. We follow a two-step approach where we first learn to predict a\nsemantic representation (SR) from video and then generate natural language\ndescriptions from the SR. To produce consistent multi-sentence descriptions, we\nmodel across-sentence consistency at the level of the SR by enforcing a\nconsistent topic. We also contribute both to the visual recognition of objects\nproposing a hand-centric approach as well as to the robust generation of\nsentences using a word lattice. Human judges rate our multi-sentence\ndescriptions as more readable, correct, and relevant than related work. To\nunderstand the difference between more detailed and shorter descriptions, we\ncollect and analyze a video description corpus of three levels of detail.","url_abs":"http://arxiv.org/abs/1403.6173v1","url_pdf":"http://arxiv.org/pdf/1403.6173v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"sentence","task_name":"Sentence"},{"task_slug":"video-description","task_name":"Video Description"}],"methods":[],"datasets_introduced":[{"slug":"tacos-multi-level-corpus","name":"TACoS Multi-Level Corpus","full_name":""}],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1403.6173","atlas_url":"https://app.syntology.ai/?focus=1403.6173","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}