{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/190408141","title":"MHP-VOS: Multiple Hypotheses Propagation for Video Object Segmentation","arxiv_id":"1904.08141","date":"2019-04-17","proceeding":"CVPR 2019 6","authors":["Shuangjie Xu","Daizong Liu","Linchao Bao","Wei Liu","Pan Zhou"],"abstract":"We address the problem of semi-supervised video object segmentation (VOS),\nwhere the masks of objects of interests are given in the first frame of an\ninput video. To deal with challenging cases where objects are occluded or\nmissing, previous work relies on greedy data association strategies that make\ndecisions for each frame individually. In this paper, we propose a novel\napproach to defer the decision making for a target object in each frame, until\na global view can be established with the entire video being taken into\nconsideration. Our approach is in the same spirit as Multiple Hypotheses\nTracking (MHT) methods, making several critical adaptations for the VOS\nproblem. We employ the bounding box (bbox) hypothesis for tracking tree\nformation, and the multiple hypotheses are spawned by propagating the preceding\nbbox into the detected bbox proposals within a gated region starting from the\ninitial object mask in the first frame. The gated region is determined by a\ngating scheme which takes into account a more comprehensive motion model rather\nthan the simple Kalman filtering model in traditional MHT. To further design\nmore customized algorithms tailored for VOS, we develop a novel mask\npropagation score instead of the appearance similarity score that could be\nbrittle due to large deformations. The mask propagation score, together with\nthe motion score, determines the affinity between the hypotheses during tree\npruning. Finally, a novel mask merging strategy is employed to handle mask\nconflicts between objects. Extensive experiments on challenging datasets\ndemonstrate the effectiveness of the proposed method, especially in the case of\nobject missing.","url_abs":"http://arxiv.org/abs/1904.08141v1","url_pdf":"http://arxiv.org/pdf/1904.08141v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"190408141","repo_url":"https://github.com/shuangjiexu/MHP-VOS","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"decision-making","task_name":"Decision Making"},{"task_slug":"object","task_name":"Object"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"},{"task_slug":"semi-supervised-video-object-segmentation","task_name":"Semi-Supervised Video Object Segmentation"},{"task_slug":"video-object-segmentation","task_name":"Video Object Segmentation"},{"task_slug":"video-semantic-segmentation","task_name":"Video Semantic Segmentation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/visual-object-tracking-on-davis-2016","task":"Semi-Supervised Video Object Segmentation","dataset":"DAVIS 2016","model":"MHP-VOS","rank_in_archive_order":40,"of":78,"metrics":{"F-measure (Decay)":"9.0","F-measure (Mean)":"89.5","F-measure (Recall)":"95.5","J&F":"88.55","Jaccard (Decay)":"6.9","Jaccard (Mean)":"87.6","Jaccard (Recall)":"97.3"},"uses_additional_data":false},{"leaderboard":"/sota/semi-supervised-video-object-segmentation-on-1","task":"Semi-Supervised Video Object Segmentation","dataset":"DAVIS 2017 (test-dev)","model":"MHP-VOS","rank_in_archive_order":41,"of":59,"metrics":{"F-measure (Decay)":"19.1","F-measure (Mean)":"72.7","F-measure (Recall)":"82.3","J&F":"69.5","Jaccard (Decay)":"18.0","Jaccard (Mean)":"66.4","Jaccard (Recall)":"76.0"},"uses_additional_data":false},{"leaderboard":"/sota/visual-object-tracking-on-davis-2017","task":"Semi-Supervised Video Object Segmentation","dataset":"DAVIS 2017 (val)","model":"MHP-VOS","rank_in_archive_order":54,"of":81,"metrics":{"F-measure (Decay)":"19.1","F-measure (Mean)":"78.9","F-measure (Recall)":"87.2","J&F":"76.15","Jaccard (Decay)":"17.8","Jaccard (Mean)":"73.4","Jaccard (Recall)":"83.5"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1904.08141","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}