{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/self-supervised-video-representation-learning-1","title":"Self-Supervised Video Representation Learning With Odd-One-Out Networks","arxiv_id":"1611.06646","date":"2016-11-21","proceeding":"CVPR 2017 7","authors":["Basura Fernando","Hakan Bilen","Efstratios Gavves","Stephen Gould"],"abstract":"We propose a new self-supervised CNN pre-training technique based on a novel\nauxiliary task called \"odd-one-out learning\". In this task, the machine is\nasked to identify the unrelated or odd element from a set of otherwise related\nelements. We apply this technique to self-supervised video representation\nlearning where we sample subsequences from videos and ask the network to learn\nto predict the odd video subsequence. The odd video subsequence is sampled such\nthat it has wrong temporal order of frames while the even ones have the correct\ntemporal order. Therefore, to generate a odd-one-out question no manual\nannotation is required. Our learning machine is implemented as multi-stream\nconvolutional neural network, which is learned end-to-end. Using odd-one-out\nnetworks, we learn temporal representations for videos that generalizes to\nother related tasks such as action recognition.\n  On action classification, our method obtains 60.3\\% on the UCF101 dataset\nusing only UCF101 data for training which is approximately 10% better than\ncurrent state-of-the-art self-supervised learning methods. Similarly, on HMDB51\ndataset we outperform self-supervised state-of-the art methods by 12.7% on\naction classification task.","url_abs":"http://arxiv.org/abs/1611.06646v4","url_pdf":"http://arxiv.org/pdf/1611.06646v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"action-classification","task_name":"Action Classification"},{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"odd-one-out","task_name":"Odd One Out"},{"task_slug":"representation-learning","task_name":"Representation Learning"},{"task_slug":"self-supervised-action-recognition","task_name":"Self-Supervised Action Recognition"},{"task_slug":"self-supervised-learning","task_name":"Self-Supervised Learning"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/self-supervised-action-recognition-on-ucf101","task":"Self-Supervised Action Recognition","dataset":"UCF101","model":"O3N (AlexNet)","rank_in_archive_order":47,"of":53,"metrics":{"3-fold Accuracy":"60.3","Frozen":"false","Pre-Training Dataset":"UCF101"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}