{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/unsupervised-learning-of-view-invariant","title":"Unsupervised Learning of View-invariant Action Representations","arxiv_id":"1809.01844","date":"2018-09-06","proceeding":"NeurIPS 2018 12","authors":["Junnan Li","Yongkang Wong","Qi Zhao","Mohan S. Kankanhalli"],"abstract":"The recent success in human action recognition with deep learning methods\nmostly adopt the supervised learning paradigm, which requires significant\namount of manually labeled data to achieve good performance. However, label\ncollection is an expensive and time-consuming process. In this work, we propose\nan unsupervised learning framework, which exploits unlabeled data to learn\nvideo representations. Different from previous works in video representation\nlearning, our unsupervised learning task is to predict 3D motion in multiple\ntarget views using video representation from a source view. By learning to\nextrapolate cross-view motions, the representation can capture view-invariant\nmotion dynamics which is discriminative for the action. In addition, we propose\na view-adversarial training method to enhance learning of view-invariant\nfeatures. We demonstrate the effectiveness of the learned representations for\naction recognition on multiple datasets.","url_abs":"http://arxiv.org/abs/1809.01844v1","url_pdf":"http://arxiv.org/pdf/1809.01844v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"unsupervised-learning-of-view-invariant","repo_url":"https://github.com/minostauros/VIAR","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"representation-learning","task_name":"Representation Learning"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1809.01844","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}