{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/video-language-modeling-a-baseline-for","title":"Video (language) modeling: a baseline for generative models of natural videos","arxiv_id":"1412.6604","date":"2014-12-20","proceeding":null,"authors":["MarcAurelio Ranzato","Arthur Szlam","Joan Bruna","Michael Mathieu","Ronan Collobert","Sumit Chopra"],"abstract":"We propose a strong baseline model for unsupervised feature learning using\nvideo data. By learning to predict missing frames or extrapolate future frames\nfrom an input video sequence, the model discovers both spatial and temporal\ncorrelations which are useful to represent complex deformations and motion\npatterns. The models we propose are largely borrowed from the language modeling\nliterature, and adapted to the vision domain by quantizing the space of image\npatches into a large dictionary. We demonstrate the approach on both a filling\nand a generation task. For the first time, we show that, after training on\nnatural videos, such a model can predict non-trivial motions over short video\nsequences.","url_abs":"http://arxiv.org/abs/1412.6604v5","url_pdf":"http://arxiv.org/pdf/1412.6604v5.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"video-language-modeling-a-baseline-for","repo_url":"https://github.com/amritanjali123/NM373_Future_Predicators","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"language-modeling","task_name":"Language Modeling"},{"task_slug":"language-modelling","task_name":"Language Modelling"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1412.6604","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}