{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/mocogan-decomposing-motion-and-content-for","title":"MoCoGAN: Decomposing Motion and Content for Video Generation","arxiv_id":"1707.04993","date":"2017-07-17","proceeding":"CVPR 2018 6","authors":["Sergey Tulyakov","Ming-Yu Liu","Xiaodong Yang","Jan Kautz"],"abstract":"Visual signals in a video can be divided into content and motion. While\ncontent specifies which objects are in the video, motion describes their\ndynamics. Based on this prior, we propose the Motion and Content decomposed\nGenerative Adversarial Network (MoCoGAN) framework for video generation. The\nproposed framework generates a video by mapping a sequence of random vectors to\na sequence of video frames. Each random vector consists of a content part and a\nmotion part. While the content part is kept fixed, the motion part is realized\nas a stochastic process. To learn motion and content decomposition in an\nunsupervised manner, we introduce a novel adversarial learning scheme utilizing\nboth image and video discriminators. Extensive experimental results on several\nchallenging datasets with qualitative and quantitative comparison to the\nstate-of-the-art approaches, verify effectiveness of the proposed framework. In\naddition, we show that MoCoGAN allows one to generate videos with same content\nbut different motion as well as videos with different content and same motion.","url_abs":"http://arxiv.org/abs/1707.04993v2","url_pdf":"http://arxiv.org/pdf/1707.04993v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"mocogan-decomposing-motion-and-content-for","repo_url":"https://github.com/sergeytulyakov/mocogan","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}},{"paper_slug":"mocogan-decomposing-motion-and-content-for","repo_url":"https://github.com/DLHacks/mocogan","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}},{"paper_slug":"mocogan-decomposing-motion-and-content-for","repo_url":"https://github.com/HappyBahman/ldvdGAN","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"mocogan-decomposing-motion-and-content-for","repo_url":"https://github.com/UBC-Computer-Vision-Group/DwNet","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"NOASSERTION"}},{"paper_slug":"mocogan-decomposing-motion-and-content-for","repo_url":"https://github.com/ubc-vision/DwNet","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"NOASSERTION"}}],"tasks":[{"task_slug":null,"task_name":"Generative Adversarial Network"},{"task_slug":"video-generation","task_name":"Video Generation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/video-generation-on-bair-robot-pushing","task":"Video Generation","dataset":"BAIR Robot Pushing","model":"MoCoGAN","rank_in_archive_order":29,"of":31,"metrics":{"Cond":"4","FVD score":"503","Pred":"12","Train":"12"},"uses_additional_data":false},{"leaderboard":"/sota/video-generation-on-ucf-101-16-frames-64x64","task":"Video Generation","dataset":"UCF-101 16 frames, 64x64, Unconditional","model":"MoCoGAN","rank_in_archive_order":4,"of":7,"metrics":{"Inception Score":"12.42"},"uses_additional_data":false},{"leaderboard":"/sota/video-generation-on-ucf-101-16-frames","task":"Video Generation","dataset":"UCF-101 16 frames, Unconditional, Single GPU","model":"MoCoGAN","rank_in_archive_order":4,"of":7,"metrics":{"Inception Score":"12.42"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1707.04993","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}