{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/actor-and-observer-joint-modeling-of-first","title":"Actor and Observer: Joint Modeling of First and Third-Person Videos","arxiv_id":"1804.09627","date":"2018-04-25","proceeding":"CVPR 2018 6","authors":["Gunnar A. Sigurdsson","Abhinav Gupta","Cordelia Schmid","Ali Farhadi","Karteek Alahari"],"abstract":"Several theories in cognitive neuroscience suggest that when people interact\nwith the world, or simulate interactions, they do so from a first-person\negocentric perspective, and seamlessly transfer knowledge between third-person\n(observer) and first-person (actor). Despite this, learning such models for\nhuman action recognition has not been achievable due to the lack of data. This\npaper takes a step in this direction, with the introduction of Charades-Ego, a\nlarge-scale dataset of paired first-person and third-person videos, involving\n112 people, with 4000 paired videos. This enables learning the link between the\ntwo, actor and observer perspectives. Thereby, we address one of the biggest\nbottlenecks facing egocentric vision research, providing a link from\nfirst-person to the abundant third-person data on the web. We use this data to\nlearn a joint representation of first and third-person videos, with only weak\nsupervision, and show its effectiveness for transferring knowledge from the\nthird-person to the first-person domain.","url_abs":"http://arxiv.org/abs/1804.09627v1","url_pdf":"http://arxiv.org/pdf/1804.09627v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"actor-and-observer-joint-modeling-of-first","repo_url":"https://github.com/gsig/actor-observer","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok","spdx":"GPL-3.0"}}],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1804.09627","atlas_url":"https://app.syntology.ai/?focus=1804.09627","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}