{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/pixel-level-matching-for-video-object","title":"Pixel-Level Matching for Video Object Segmentation using Convolutional Neural Networks","arxiv_id":"1708.05137","date":"2017-08-17","proceeding":"ICCV 2017 10","authors":["Jae Shin Yoon","Francois Rameau","Junsik Kim","Seokju Lee","Seunghak Shin","In So Kweon"],"abstract":"We propose a novel video object segmentation algorithm based on pixel-level\nmatching using Convolutional Neural Networks (CNN). Our network aims to\ndistinguish the target area from the background on the basis of the pixel-level\nsimilarity between two object units. The proposed network represents a target\nobject using features from different depth layers in order to take advantage of\nboth the spatial details and the category-level semantic information.\nFurthermore, we propose a feature compression technique that drastically\nreduces the memory requirements while maintaining the capability of feature\nrepresentation. Two-stage training (pre-training and fine-tuning) allows our\nnetwork to handle any target object regardless of its category (even if the\nobject's type does not belong to the pre-training data) or of variations in its\nappearance through a video sequence. Experiments on large datasets demonstrate\nthe effectiveness of our model - against related methods - in terms of\naccuracy, speed, and stability. Finally, we introduce the transferability of\nour network to different domains, such as the infrared data domain.","url_abs":"http://arxiv.org/abs/1708.05137v1","url_pdf":"http://arxiv.org/pdf/1708.05137v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"feature-compression","task_name":"Feature Compression"},{"task_slug":"object","task_name":"Object"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"},{"task_slug":"semi-supervised-video-object-segmentation","task_name":"Semi-Supervised Video Object Segmentation"},{"task_slug":"video-object-segmentation","task_name":"Video Object Segmentation"},{"task_slug":"video-semantic-segmentation","task_name":"Video Semantic Segmentation"},{"task_slug":"visual-object-tracking","task_name":"Visual Object Tracking"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/visual-object-tracking-on-davis-2016","task":"Semi-Supervised Video Object Segmentation","dataset":"DAVIS 2016","model":"PLM","rank_in_archive_order":73,"of":78,"metrics":{"F-measure (Decay)":"14.7","F-measure (Mean)":"62.5","F-measure (Recall)":"73.2","J&F":"66.35","Jaccard (Decay)":"11.2","Jaccard (Mean)":"70.2","Jaccard (Recall)":"86.3"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}