{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/joint-person-segmentation-and-identification","title":"Joint Person Segmentation and Identification in Synchronized First- and Third-person Videos","arxiv_id":"1803.11217","date":"2018-03-29","proceeding":"ECCV 2018 9","authors":["Mingze Xu","Chenyou Fan","Yuchen Wang","Michael S. Ryoo","David J. Crandall"],"abstract":"In a world of pervasive cameras, public spaces are often captured from\nmultiple perspectives by cameras of different types, both fixed and mobile. An\nimportant problem is to organize these heterogeneous collections of videos by\nfinding connections between them, such as identifying correspondences between\nthe people appearing in the videos and the people holding or wearing the\ncameras. In this paper, we wish to solve two specific problems: (1) given two\nor more synchronized third-person videos of a scene, produce a pixel-level\nsegmentation of each visible person and identify corresponding people across\ndifferent views (i.e., determine who in camera A corresponds with whom in\ncamera B), and (2) given one or more synchronized third-person videos as well\nas a first-person video taken by a mobile or wearable camera, segment and\nidentify the camera wearer in the third-person videos. Unlike previous work\nwhich requires ground truth bounding boxes to estimate the correspondences, we\nperform person segmentation and identification jointly. We find that solving\nthese two problems simultaneously is mutually beneficial, because better\nfine-grained segmentation allows us to better perform matching across views,\nand information from multiple views helps us perform more accurate\nsegmentation. We evaluate our approach on two challenging datasets of\ninteracting people captured from multiple wearable cameras, and show that our\nproposed method performs significantly better than the state-of-the-art on both\nperson segmentation and identification.","url_abs":"http://arxiv.org/abs/1803.11217v2","url_pdf":"http://arxiv.org/pdf/1803.11217v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"segmentation","task_name":"Segmentation"}],"methods":[],"datasets_introduced":[{"slug":"iu-shareview","name":"IU ShareView","full_name":""}],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1803.11217","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}