{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/time-contrastive-networks-self-supervised","title":"Time-Contrastive Networks: Self-Supervised Learning from Video","arxiv_id":"1704.06888","date":"2017-04-23","proceeding":null,"authors":["Pierre Sermanet","Corey Lynch","Yevgen Chebotar","Jasmine Hsu","Eric Jang","Stefan Schaal","Sergey Levine"],"abstract":"We propose a self-supervised approach for learning representations and\nrobotic behaviors entirely from unlabeled videos recorded from multiple\nviewpoints, and study how this representation can be used in two robotic\nimitation settings: imitating object interactions from videos of humans, and\nimitating human poses. Imitation of human behavior requires a\nviewpoint-invariant representation that captures the relationships between\nend-effectors (hands or robot grippers) and the environment, object attributes,\nand body pose. We train our representations using a metric learning loss, where\nmultiple simultaneous viewpoints of the same observation are attracted in the\nembedding space, while being repelled from temporal neighbors which are often\nvisually similar but functionally different. In other words, the model\nsimultaneously learns to recognize what is common between different-looking\nimages, and what is different between similar-looking images. This signal\ncauses our model to discover attributes that do not change across viewpoint,\nbut do change across time, while ignoring nuisance variables such as\nocclusions, motion blur, lighting and background. We demonstrate that this\nrepresentation can be used by a robot to directly mimic human poses without an\nexplicit correspondence, and that it can be used as a reward function within a\nreinforcement learning algorithm. While representations are learned from an\nunlabeled collection of task-related videos, robot behaviors such as pouring\nare learned by watching a single 3rd-person demonstration by a human. Reward\nfunctions obtained by following the human demonstrations under the learned\nrepresentation enable efficient reinforcement learning that is practical for\nreal-world robotic systems. Video results, open-source code and dataset are\navailable at https://sermanet.github.io/imitate","url_abs":"http://arxiv.org/abs/1704.06888v3","url_pdf":"http://arxiv.org/pdf/1704.06888v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"time-contrastive-networks-self-supervised","repo_url":"https://github.com/araffin/srl-zoo","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}},{"paper_slug":"time-contrastive-networks-self-supervised","repo_url":"https://github.com/elicassion/3dtrl","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"time-contrastive-networks-self-supervised","repo_url":"https://github.com/sermanet/tcn","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"unanswered"}},{"paper_slug":"time-contrastive-networks-self-supervised","repo_url":"https://github.com/tensorflow/models","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null},{"paper_slug":"time-contrastive-networks-self-supervised","repo_url":"https://github.com/tensorflow/models/tree/master/research/tcn","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null},{"paper_slug":"time-contrastive-networks-self-supervised","repo_url":"https://github.com/2023-MindSpore-1/ms-code-7/tree/main/TCN","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null},{"paper_slug":"time-contrastive-networks-self-supervised","repo_url":"https://github.com/tensorflow/models/tree/archive/research/tcn","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"tf","reach":null}],"tasks":[{"task_slug":"metric-learning","task_name":"Metric Learning"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"self-supervised-learning","task_name":"Self-Supervised Learning"},{"task_slug":"video-alignment","task_name":"Video Alignment"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[{"method_slug":"1cycle","method_name":"1cycle"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/video-alignment-on-upenn-action","task":"Video Alignment","dataset":"UPenn Action","model":"TCN","rank_in_archive_order":3,"of":4,"metrics":{"Kendall's Tau":"0.7353"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1704.06888","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}