{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/cnn-in-mrf-video-object-segmentation-via","title":"CNN in MRF: Video Object Segmentation via Inference in A CNN-Based Higher-Order Spatio-Temporal MRF","arxiv_id":"1803.09453","date":"2018-03-26","proceeding":"CVPR 2018 6","authors":["Linchao Bao","Baoyuan Wu","Wei Liu"],"abstract":"This paper addresses the problem of video object segmentation, where the\ninitial object mask is given in the first frame of an input video. We propose a\nnovel spatio-temporal Markov Random Field (MRF) model defined over pixels to\nhandle this problem. Unlike conventional MRF models, the spatial dependencies\namong pixels in our model are encoded by a Convolutional Neural Network (CNN).\nSpecifically, for a given object, the probability of a labeling to a set of\nspatially neighboring pixels can be predicted by a CNN trained for this\nspecific object. As a result, higher-order, richer dependencies among pixels in\nthe set can be implicitly modeled by the CNN. With temporal dependencies\nestablished by optical flow, the resulting MRF model combines both spatial and\ntemporal cues for tackling video object segmentation. However, performing\ninference in the MRF model is very difficult due to the very high-order\ndependencies. To this end, we propose a novel CNN-embedded algorithm to perform\napproximate inference in the MRF. This algorithm proceeds by alternating\nbetween a temporal fusion step and a feed-forward CNN step. When initialized\nwith an appearance-based one-shot segmentation CNN, our model outperforms the\nwinning entries of the DAVIS 2017 Challenge, without resorting to model\nensembling or any dedicated detectors.","url_abs":"http://arxiv.org/abs/1803.09453v1","url_pdf":"http://arxiv.org/pdf/1803.09453v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"object","task_name":"Object"},{"task_slug":"one-shot-segmentation","task_name":"One-Shot Segmentation"},{"task_slug":"optical-flow-estimation","task_name":"Optical Flow Estimation"},{"task_slug":"segmentation","task_name":"Segmentation"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"},{"task_slug":"semi-supervised-video-object-segmentation","task_name":"Semi-Supervised Video Object Segmentation"},{"task_slug":"video-object-segmentation","task_name":"Video Object Segmentation"},{"task_slug":"video-semantic-segmentation","task_name":"Video Semantic Segmentation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/visual-object-tracking-on-davis-2016","task":"Semi-Supervised Video Object Segmentation","dataset":"DAVIS 2016","model":"CINM","rank_in_archive_order":52,"of":78,"metrics":{"F-measure (Decay)":"14.7","F-measure (Mean)":"85.0","F-measure (Recall)":"92.1","J&F":"84.2","Jaccard (Decay)":"12.3","Jaccard (Mean)":"83.4","Jaccard (Recall)":"94.9"},"uses_additional_data":false},{"leaderboard":"/sota/semi-supervised-video-object-segmentation-on-1","task":"Semi-Supervised Video Object Segmentation","dataset":"DAVIS 2017 (test-dev)","model":"CINM","rank_in_archive_order":42,"of":59,"metrics":{"F-measure (Decay)":"20.0","F-measure (Mean)":"70.5","F-measure (Recall)":"79.6","J&F":"67.5","Jaccard (Decay)":"20.0","Jaccard (Mean)":"64.5","Jaccard (Recall)":"73.8"},"uses_additional_data":false},{"leaderboard":"/sota/visual-object-tracking-on-davis-2017","task":"Semi-Supervised Video Object Segmentation","dataset":"DAVIS 2017 (val)","model":"CINM","rank_in_archive_order":59,"of":81,"metrics":{"F-measure (Decay)":"26.2","F-measure (Mean)":"74.0","F-measure (Recall)":"81.6","J&F":"70.6","Jaccard (Decay)":"24.6","Jaccard (Mean)":"67.2","Jaccard (Recall)":"74.5"},"uses_additional_data":false},{"leaderboard":"/sota/video-object-segmentation-on-youtube","task":"Semi-Supervised Video Object Segmentation","dataset":"YouTube","model":"MRFCNN","rank_in_archive_order":2,"of":5,"metrics":{"mIoU":"0.784"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1803.09453","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}