{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/self-supervised-spatio-temporal","title":"Self-supervised Spatio-temporal Representation Learning for Videos by Predicting Motion and Appearance Statistics","arxiv_id":"1904.03597","date":"2019-04-07","proceeding":"CVPR 2019 6","authors":["Jiangliu Wang","Jianbo Jiao","Linchao Bao","Shengfeng He","Yun-hui Liu","Wei Liu"],"abstract":"We address the problem of video representation learning without\nhuman-annotated labels. While previous efforts address the problem by designing\nnovel self-supervised tasks using video data, the learned features are merely\non a frame-by-frame basis, which are not applicable to many video analytic\ntasks where spatio-temporal features are prevailing. In this paper we propose a\nnovel self-supervised approach to learn spatio-temporal features for video\nrepresentation. Inspired by the success of two-stream approaches in video\nclassification, we propose to learn visual features by regressing both motion\nand appearance statistics along spatial and temporal dimensions, given only the\ninput video data. Specifically, we extract statistical concepts (fast-motion\nregion and the corresponding dominant direction, spatio-temporal color\ndiversity, dominant color, etc.) from simple patterns in both spatial and\ntemporal domains. Unlike prior puzzles that are even hard for humans to solve,\nthe proposed approach is consistent with human inherent visual habits and\ntherefore easy to answer. We conduct extensive experiments with C3D to validate\nthe effectiveness of our proposed approach. The experiments show that our\napproach can significantly improve the performance of C3D when applied to video\nclassification tasks. Code is available at\nhttps://github.com/laura-wang/video_repres_mas.","url_abs":"http://arxiv.org/abs/1904.03597v1","url_pdf":"http://arxiv.org/pdf/1904.03597v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"self-supervised-spatio-temporal","repo_url":"https://github.com/laura-wang/video_repres_mas","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"tf","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"representation-learning","task_name":"Representation Learning"},{"task_slug":"self-supervised-action-recognition","task_name":"Self-Supervised Action Recognition"},{"task_slug":"video-classification","task_name":"Video Classification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/self-supervised-action-recognition-on-hmdb51","task":"Self-Supervised Action Recognition","dataset":"HMDB51","model":"Motion & Appearance (C3D)","rank_in_archive_order":47,"of":48,"metrics":{"Frozen":"false","Pre-Training Dataset":"UCF101","Top-1 Accuracy":"20.3"},"uses_additional_data":false},{"leaderboard":"/sota/self-supervised-action-recognition-on-ucf101","task":"Self-Supervised Action Recognition","dataset":"UCF101","model":"Motion & Appearance (C3D)","rank_in_archive_order":49,"of":53,"metrics":{"3-fold Accuracy":"58.8","Frozen":"false","Pre-Training Dataset":"UCF101"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1904.03597","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}