{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/video-generation-from-single-semantic-label","title":"Video Generation from Single Semantic Label Map","arxiv_id":"1903.04480","date":"2019-03-11","proceeding":"CVPR 2019 6","authors":["Junting Pan","Chengyu Wang","Xu Jia","Jing Shao","Lu Sheng","Junjie Yan","Xiaogang Wang"],"abstract":"This paper proposes the novel task of video generation conditioned on a\nSINGLE semantic label map, which provides a good balance between flexibility\nand quality in the generation process. Different from typical end-to-end\napproaches, which model both scene content and dynamics in a single step, we\npropose to decompose this difficult task into two sub-problems. As current\nimage generation methods do better than video generation in terms of detail, we\nsynthesize high quality content by only generating the first frame. Then we\nanimate the scene based on its semantic meaning to obtain the temporally\ncoherent video, giving us excellent results overall. We employ a cVAE for\npredicting optical flow as a beneficial intermediate step to generate a video\nsequence conditioned on the initial single frame. A semantic label map is\nintegrated into the flow prediction module to achieve major improvements in the\nimage-to-video generation process. Extensive experiments on the Cityscapes\ndataset show that our method outperforms all competing methods.","url_abs":"http://arxiv.org/abs/1903.04480v1","url_pdf":"http://arxiv.org/pdf/1903.04480v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"video-generation-from-single-semantic-label","repo_url":"https://github.com/junting/seg2vid","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"video-generation-from-single-semantic-label","repo_url":"https://github.com/STVIR/seg2vid","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"image-generation","task_name":"Image Generation"},{"task_slug":"image-to-video","task_name":"Image to Video Generation"},{"task_slug":"optical-flow-estimation","task_name":"Optical Flow Estimation"},{"task_slug":"video-generation","task_name":"Video Generation"}],"methods":[{"method_slug":"cvae","method_name":"cVAE"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1903.04480","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}