{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/hierarchical-video-generation-from-orthogonal","title":"Hierarchical Video Generation from Orthogonal Information: Optical Flow and Texture","arxiv_id":"1711.09618","date":"2017-11-27","proceeding":null,"authors":["Katsunori Ohnishi","Shohei Yamamoto","Yoshitaka Ushiku","Tatsuya Harada"],"abstract":"Learning to represent and generate videos from unlabeled data is a very\nchallenging problem. To generate realistic videos, it is important not only to\nensure that the appearance of each frame is real, but also to ensure the\nplausibility of a video motion and consistency of a video appearance in the\ntime direction. The process of video generation should be divided according to\nthese intrinsic difficulties. In this study, we focus on the motion and\nappearance information as two important orthogonal components of a video, and\npropose Flow-and-Texture-Generative Adversarial Networks (FTGAN) consisting of\nFlowGAN and TextureGAN. In order to avoid a huge annotation cost, we have to\nexplore a way to learn from unlabeled data. Thus, we employ optical flow as\nmotion information to generate videos. FlowGAN generates optical flow, which\ncontains only the edge and motion of the videos to be begerated. On the other\nhand, TextureGAN specializes in giving a texture to optical flow generated by\nFlowGAN. This hierarchical approach brings more realistic videos with plausible\nmotion and appearance consistency. Our experiments show that our model\ngenerates more plausible motion videos and also achieves significantly improved\nperformance for unsupervised action classification in comparison to previous\nGAN works. In addition, because our model generates videos from two independent\ninformation, our model can generate new combinations of motion and attribute\nthat are not seen in training data, such as a video in which a person is doing\nsit-up in a baseball ground.","url_abs":"http://arxiv.org/abs/1711.09618v2","url_pdf":"http://arxiv.org/pdf/1711.09618v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"hierarchical-video-generation-from-orthogonal","repo_url":"https://github.com/mil-tokyo/FTGAN","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"unanswered"}},{"paper_slug":"hierarchical-video-generation-from-orthogonal","repo_url":"https://github.com/synce1234/FTGAN_custom","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"unanswered"}},{"paper_slug":"hierarchical-video-generation-from-orthogonal","repo_url":"https://github.com/vikramjit-sidhu/hlcv_project","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"action-classification","task_name":"Action Classification"},{"task_slug":"attribute","task_name":"Attribute"},{"task_slug":"optical-flow-estimation","task_name":"Optical Flow Estimation"},{"task_slug":"video-generation","task_name":"Video Generation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1711.09618","atlas_url":"https://app.syntology.ai/?focus=1711.09618","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}