{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/from-third-person-to-first-person-dataset-and","title":"From Third Person to First Person: Dataset and Baselines for Synthesis and Retrieval","arxiv_id":"1812.00104","date":"2018-12-01","proceeding":null,"authors":["Mohamed Elfeki","Krishna Regmi","Shervin Ardeshir","Ali Borji"],"abstract":"First-person (egocentric) and third person (exocentric) videos are\ndrastically different in nature. The relationship between these two views have\nbeen studied in recent years, however, it has yet to be fully explored. In this\nwork, we introduce two datasets (synthetic and natural/real) containing\nsimultaneously recorded egocentric and exocentric videos. We also explore\nrelating the two domains (egocentric and exocentric) in two aspects. First, we\nsynthesize images in the egocentric domain from the exocentric domain using a\nconditional generative adversarial network (cGAN). We show that with enough\ntraining data, our network is capable of hallucinating how the world would look\nlike from an egocentric perspective, given an exocentric video. Second, we\naddress the cross-view retrieval problem across the two views. Given an\negocentric query frame (or its momentary optical flow), we retrieve its\ncorresponding exocentric frame (or optical flow) from a gallery set. We show\nthat using synthetic data could be beneficial in retrieving real data. We show\nthat performing domain adaptation from the synthetic domain to the natural/real\ndomain, is helpful in tasks such as retrieval. We believe that the presented\ndatasets and the proposed baselines offer new opportunities for further\nresearch in this direction. The code and dataset are publicly available.","url_abs":"http://arxiv.org/abs/1812.00104v1","url_pdf":"http://arxiv.org/pdf/1812.00104v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"from-third-person-to-first-person-dataset-and","repo_url":"https://github.com/M-Elfeki/ThirdToFirst","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"domain-adaptation","task_name":"Domain Adaptation"},{"task_slug":null,"task_name":"Generative Adversarial Network"},{"task_slug":"optical-flow-estimation","task_name":"Optical Flow Estimation"},{"task_slug":"retrieval","task_name":"Retrieval"}],"methods":[],"datasets_introduced":[{"slug":"thirdtofirst","name":"ThirdToFirst","full_name":""}],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1812.00104","atlas_url":"https://app.syntology.ai/?focus=1812.00104","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}