{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/probabilistic-video-generation-using-holistic","title":"Probabilistic Video Generation using Holistic Attribute Control","arxiv_id":"1803.08085","date":"2018-03-21","proceeding":"ECCV 2018 9","authors":["Jiawei He","Andreas Lehrmann","Joseph Marino","Greg Mori","Leonid Sigal"],"abstract":"Videos express highly structured spatio-temporal patterns of visual data. A\nvideo can be thought of as being governed by two factors: (i) temporally\ninvariant (e.g., person identity), or slowly varying (e.g., activity),\nattribute-induced appearance, encoding the persistent content of each frame,\nand (ii) an inter-frame motion or scene dynamics (e.g., encoding evolution of\nthe person ex-ecuting the action). Based on this intuition, we propose a\ngenerative framework for video generation and future prediction. The proposed\nframework generates a video (short clip) by decoding samples sequentially drawn\nfrom a latent space distribution into full video frames. Variational\nAutoencoders (VAEs) are used as a means of encoding/decoding frames into/from\nthe latent space and RNN as a wayto model the dynamics in the latent space. We\nimprove the video generation consistency through temporally-conditional\nsampling and quality by structuring the latent space with attribute controls;\nensuring that attributes can be both inferred and conditioned on during\nlearning/generation. As a result, given attributes and/orthe first frame, our\nmodel is able to generate diverse but highly consistent sets ofvideo sequences,\naccounting for the inherent uncertainty in the prediction task. Experimental\nresults on Chair CAD, Weizmann Human Action, and MIT-Flickr datasets, along\nwith detailed comparison to the state-of-the-art, verify effectiveness of the\nframework.","url_abs":"http://arxiv.org/abs/1803.08085v1","url_pdf":"http://arxiv.org/pdf/1803.08085v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"probabilistic-video-generation-using-holistic","repo_url":"https://github.com/charlescheng0117/pytorch-VideoVAE","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"attribute","task_name":"Attribute"},{"task_slug":"future-prediction","task_name":"Future prediction"},{"task_slug":"video-generation","task_name":"Video Generation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1803.08085","atlas_url":"https://app.syntology.ai/?focus=1803.08085","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}