{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/analysis-of-hand-segmentation-in-the-wild","title":"Analysis of Hand Segmentation in the Wild","arxiv_id":"1803.03317","date":"2018-03-08","proceeding":"CVPR 2018 6","authors":["Aisha Urooj Khan","Ali Borji"],"abstract":"A large number of works in egocentric vision have concentrated on action and\nobject recognition. Detection and segmentation of hands in first-person videos,\nhowever, has less been explored. For many applications in this domain, it is\nnecessary to accurately segment not only hands of the camera wearer but also\nthe hands of others with whom he is interacting. Here, we take an in-depth look\nat the hand segmentation problem. In the quest for robust hand segmentation\nmethods, we evaluated the performance of the state of the art semantic\nsegmentation methods, off the shelf and fine-tuned, on existing datasets. We\nfine-tune RefineNet, a leading semantic segmentation method, for hand\nsegmentation and find that it does much better than the best contenders.\nExisting hand segmentation datasets are collected in the laboratory settings.\nTo overcome this limitation, we contribute by collecting two new datasets: a)\nEgoYouTubeHands including egocentric videos containing hands in the wild, and\nb) HandOverFace to analyze the performance of our models in presence of similar\nappearance occlusions. We further explore whether conditional random fields can\nhelp refine generated hand segmentations. To demonstrate the benefit of\naccurate hand maps, we train a CNN for hand-based activity recognition and\nachieve higher accuracy when a CNN was trained using hand maps produced by the\nfine-tuned RefineNet. Finally, we annotate a subset of the EgoHands dataset for\nfine-grained action recognition and show that an accuracy of 58.6% can be\nachieved by just looking at a single hand pose which is much better than the\nchance level (12.5%).","url_abs":"http://arxiv.org/abs/1803.03317v2","url_pdf":"http://arxiv.org/pdf/1803.03317v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"analysis-of-hand-segmentation-in-the-wild","repo_url":"https://github.com/aurooj/Hand-Segmentation-in-the-Wild","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"activity-recognition","task_name":"Activity Recognition"},{"task_slug":"fine-grained-action-recognition","task_name":"Fine-grained Action Recognition"},{"task_slug":"hand-segmentation","task_name":"Hand Segmentation"},{"task_slug":"segmentation","task_name":"Segmentation"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"}],"methods":[],"datasets_introduced":[{"slug":"eyth","name":"EYTH","full_name":"EgoYouTubeHands"}],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1803.03317","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}