{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/unsupervised-learning-from-continuous-video","title":"Unsupervised Learning from Continuous Video in a Scalable Predictive Recurrent Network","arxiv_id":"1607.06854","date":"2016-07-22","proceeding":null,"authors":["Filip Piekniewski","Patryk Laurent","Csaba Petre","Micah Richert","Dimitry Fisher","Todd Hylton"],"abstract":"Understanding visual reality involves acquiring common-sense knowledge about\ncountless regularities in the visual world, e.g., how illumination alters the\nappearance of objects in a scene, and how motion changes their apparent spatial\nrelationship. These regularities are hard to label for training supervised\nmachine learning algorithms; consequently, algorithms need to learn these\nregularities from the real world in an unsupervised way. We present a novel\nnetwork meta-architecture that can learn world dynamics from raw, continuous\nvideo. The components of this network can be implemented using any algorithm\nthat possesses three key capabilities: prediction of a signal over time,\nreduction of signal dimensionality (compression), and the ability to use\nsupplementary contextual information to inform the prediction. The presented\narchitecture is highly-parallelized and scalable, and is implemented using\nlocalized connectivity, processing, and learning. We demonstrate an\nimplementation of this architecture where the components are built from\nmulti-layer perceptrons. We apply the implementation to create a system capable\nof stable and robust visual tracking of objects as seen by a moving camera.\nResults show performance on par with or exceeding state-of-the-art tracking\nalgorithms. The tracker can be trained in either fully supervised or\nunsupervised-then-briefly-supervised regimes. Success of the briefly-supervised\nregime suggests that the unsupervised portion of the model extracts useful\ninformation about visual reality. The results suggest a new class of AI\nalgorithms that uniquely combine prediction and scalability in a way that makes\nthem suitable for learning from and --- and eventually acting within --- the\nreal world.","url_abs":"http://arxiv.org/abs/1607.06854v3","url_pdf":"http://arxiv.org/pdf/1607.06854v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"unsupervised-learning-from-continuous-video","repo_url":"https://github.com/braincorp/PVM","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"NOASSERTION"}},{"paper_slug":"unsupervised-learning-from-continuous-video","repo_url":"https://github.com/mhazoglou/PVM_PyCUDA","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"common-sense-reasoning","task_name":"Common Sense Reasoning"},{"task_slug":"visual-tracking","task_name":"Visual Tracking"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1607.06854","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}