{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/unsupervised-learning-of-important-objects","title":"Unsupervised Learning of Important Objects from First-Person Videos","arxiv_id":"1611.05335","date":"2016-11-16","proceeding":"ICCV 2017 10","authors":["Gedas Bertasius","Hyun Soo Park","Stella X. Yu","Jianbo Shi"],"abstract":"A first-person camera, placed at a person's head, captures, which objects are\nimportant to the camera wearer. Most prior methods for this task learn to\ndetect such important objects from the manually labeled first-person data in a\nsupervised fashion. However, important objects are strongly related to the\ncamera wearer's internal state such as his intentions and attention, and thus,\nonly the person wearing the camera can provide the importance labels. Such a\nconstraint makes the annotation process costly and limited in scalability.\n  In this work, we show that we can detect important objects in first-person\nimages without the supervision by the camera wearer or even third-person\nlabelers. We formulate an important detection problem as an interplay between\nthe 1) segmentation and 2) recognition agents. The segmentation agent first\nproposes a possible important object segmentation mask for each image, and then\nfeeds it to the recognition agent, which learns to predict an important object\nmask using visual semantics and spatial features.\n  We implement such an interplay between both agents via an alternating\ncross-pathway supervision scheme inside our proposed Visual-Spatial Network\n(VSN). Our VSN consists of spatial (\"where\") and visual (\"what\") pathways, one\nof which learns common visual semantics while the other focuses on the spatial\nlocation cues. Our unsupervised learning is accomplished via a cross-pathway\nsupervision, where one pathway feeds its predictions to a segmentation agent,\nwhich proposes a candidate important object segmentation mask that is then used\nby the other pathway as a supervisory signal. We show our method's success on\ntwo different important object datasets, where our method achieves similar or\nbetter results as the supervised methods.","url_abs":"http://arxiv.org/abs/1611.05335v3","url_pdf":"http://arxiv.org/pdf/1611.05335v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"unsupervised-learning-of-important-objects","repo_url":"https://github.com/gberta/Visual-Spatial-Network","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"object","task_name":"Object"},{"task_slug":"segmentation","task_name":"Segmentation"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}