{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/attentive-semantic-video-generation-using","title":"Attentive Semantic Video Generation using Captions","arxiv_id":"1708.05980","date":"2017-08-20","proceeding":"ICCV 2017 10","authors":["Tanya Marwah","Gaurav Mittal","Vineeth N. Balasubramanian"],"abstract":"This paper proposes a network architecture to perform variable length\nsemantic video generation using captions. We adopt a new perspective towards\nvideo generation where we allow the captions to be combined with the long-term\nand short-term dependencies between video frames and thus generate a video in\nan incremental manner. Our experiments demonstrate our network architecture's\nability to distinguish between objects, actions and interactions in a video and\ncombine them to generate videos for unseen captions. The network also exhibits\nthe capability to perform spatio-temporal style transfer when asked to generate\nvideos for a sequence of captions. We also show that the network's ability to\nlearn a latent representation allows it generate videos in an unsupervised\nmanner and perform other tasks such as action recognition. (Accepted in\nInternational Conference in Computer Vision (ICCV) 2017)","url_abs":"http://arxiv.org/abs/1708.05980v3","url_pdf":"http://arxiv.org/pdf/1708.05980v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"attentive-semantic-video-generation-using","repo_url":"https://github.com/Singularity42/cap2vid","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"style-transfer","task_name":"Style Transfer"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"},{"task_slug":"video-generation","task_name":"Video Generation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1708.05980","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}