{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/adversarial-inference-for-multi-sentence","title":"Adversarial Inference for Multi-Sentence Video Description","arxiv_id":"1812.05634","date":"2018-12-13","proceeding":"CVPR 2019 6","authors":["Jae Sung Park","Marcus Rohrbach","Trevor Darrell","Anna Rohrbach"],"abstract":"While significant progress has been made in the image captioning task, video\ndescription is still in its infancy due to the complex nature of video data.\nGenerating multi-sentence descriptions for long videos is even more\nchallenging. Among the main issues are the fluency and coherence of the\ngenerated descriptions, and their relevance to the video. Recently,\nreinforcement and adversarial learning based methods have been explored to\nimprove the image captioning models; however, both types of methods suffer from\na number of issues, e.g. poor readability and high redundancy for RL and\nstability issues for GANs. In this work, we instead propose to apply\nadversarial techniques during inference, designing a discriminator which\nencourages better multi-sentence video description. In addition, we find that a\nmulti-discriminator \"hybrid\" design, where each discriminator targets one\naspect of a description, leads to the best results. Specifically, we decouple\nthe discriminator to evaluate on three criteria: 1) visual relevance to the\nvideo, 2) language diversity and fluency, and 3) coherence across sentences.\nOur approach results in more accurate, diverse, and coherent multi-sentence\nvideo descriptions, as shown by automatic as well as human evaluation on the\npopular ActivityNet Captions dataset.","url_abs":"http://arxiv.org/abs/1812.05634v2","url_pdf":"http://arxiv.org/pdf/1812.05634v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"adversarial-inference-for-multi-sentence","repo_url":"https://github.com/jamespark3922/adv-inf","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"diversity","task_name":"Diversity"},{"task_slug":"image-captioning","task_name":"Image Captioning"},{"task_slug":"sentence","task_name":"Sentence"},{"task_slug":"video-description","task_name":"Video Description"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1812.05634","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}