{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/unsupervised-learning-from-videos-using","title":"Unsupervised learning from videos using temporal coherency deep networks","arxiv_id":"1801.08100","date":"2018-01-24","proceeding":null,"authors":["Carolina Redondo-Cabrera","Roberto J. López-Sastre"],"abstract":"In this work we address the challenging problem of unsupervised learning from\nvideos. Existing methods utilize the spatio-temporal continuity in contiguous\nvideo frames as regularization for the learning process. Typically, this\ntemporal coherence of close frames is used as a free form of annotation,\nencouraging the learned representations to exhibit small differences between\nthese frames. But this type of approach fails to capture the dissimilarity\nbetween videos with different content, hence learning less discriminative\nfeatures. We here propose two Siamese architectures for Convolutional Neural\nNetworks, and their corresponding novel loss functions, to learn from unlabeled\nvideos, which jointly exploit the local temporal coherence between contiguous\nframes, and a global discriminative margin used to separate representations of\ndifferent videos. An extensive experimental evaluation is presented, where we\nvalidate the proposed models on various tasks. First, we show how the learned\nfeatures can be used to discover actions and scenes in video collections.\nSecond, we show the benefits of such an unsupervised learning from just\nunlabeled videos, which can be directly used as a prior for the supervised\nrecognition tasks of actions and objects in images, where our results further\nshow that our features can even surpass a traditional and heavily supervised\npre-training plus fine-tunning strategy.","url_abs":"http://arxiv.org/abs/1801.08100v2","url_pdf":"http://arxiv.org/pdf/1801.08100v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"unsupervised-learning-from-videos-using","repo_url":"https://github.com/gramuah/unsupervised","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"caffe2","reach":null}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}