{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/towards-robust-video-object-segmentation-with","title":"Towards Robust Video Object Segmentation with Adaptive Object Calibration","arxiv_id":"2207.00887","date":"2022-07-02","proceeding":null,"authors":["Xiaohao Xu","Jinglu Wang","Xiang Ming","Yan Lu"],"abstract":"In the booming video era, video segmentation attracts increasing research attention in the multimedia community. Semi-supervised video object segmentation (VOS) aims at segmenting objects in all target frames of a video, given annotated object masks of reference frames. Most existing methods build pixel-wise reference-target correlations and then perform pixel-wise tracking to obtain target masks. Due to neglecting object-level cues, pixel-level approaches make the tracking vulnerable to perturbations, and even indiscriminate among similar objects. Towards robust VOS, the key insight is to calibrate the representation and mask of each specific object to be expressive and discriminative. Accordingly, we propose a new deep network, which can adaptively construct object representations and calibrate object masks to achieve stronger robustness. First, we construct the object representations by applying an adaptive object proxy (AOP) aggregation method, where the proxies represent arbitrary-shaped segments at multi-levels for reference. Then, prototype masks are initially generated from the reference-target correlations based on AOP. Afterwards, such proto-masks are further calibrated through network modulation, conditioning on the object proxy representations. We consolidate this conditional mask calibration process in a progressive manner, where the object representations and proto-masks evolve to be discriminative iteratively. Extensive experiments are conducted on the standard VOS benchmarks, YouTube-VOS-18/19 and DAVIS-17. Our model achieves the state-of-the-art performance among existing published works, and also exhibits superior robustness against perturbations. Our project repo is at https://github.com/JerryX1110/Robust-Video-Object-Segmentation","url_abs":"https://arxiv.org/abs/2207.00887v1","url_pdf":"https://arxiv.org/pdf/2207.00887v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"towards-robust-video-object-segmentation-with","repo_url":"https://github.com/jerryx1110/robust-video-object-segmentation","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"object","task_name":"Object"},{"task_slug":"segmentation","task_name":"Segmentation"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"},{"task_slug":"semi-supervised-video-object-segmentation","task_name":"Semi-Supervised Video Object Segmentation"},{"task_slug":"video-object-segmentation","task_name":"Video Object Segmentation"},{"task_slug":"video-segmentation","task_name":"Video Segmentation"},{"task_slug":"video-semantic-segmentation","task_name":"Video Semantic Segmentation"},{"task_slug":"visual-object-tracking","task_name":"Visual Object Tracking"}],"methods":[{"method_slug":"vos","method_name":"VOS"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/video-object-segmentation-on-davis-2016","task":"Video Object Segmentation","dataset":"DAVIS 2016","model":"AOC-MF (val)","rank_in_archive_order":16,"of":24,"metrics":{"F-Score":"94.7","Jaccard (Mean)":"88.5"},"uses_additional_data":false},{"leaderboard":"/sota/video-object-segmentation-on-davis-2017","task":"Video Object Segmentation","dataset":"DAVIS 2017","model":"AOC-MF (val)","rank_in_archive_order":1,"of":5,"metrics":{"F-Score":"85.9","Jaccard (Mean)":"81.7"},"uses_additional_data":false},{"leaderboard":"/sota/visual-object-tracking-on-youtube-vos-1","task":"Visual Object Tracking","dataset":"YouTube-VOS","model":"AOC-MF","rank_in_archive_order":1,"of":2,"metrics":{"F-Measure (Seen)":"87.4","F-Measure (Unseen)":"87.1","Jaccard (Seen)":"82.7","Jaccard (Unseen)":"78.8","O (Average of Measures)":"84"},"uses_additional_data":false},{"leaderboard":"/sota/visual-object-tracking-on-youtube-vos-1","task":"Visual Object Tracking","dataset":"YouTube-VOS","model":"AOC-Base","rank_in_archive_order":2,"of":2,"metrics":{"F-Measure (Seen)":"87.2","F-Measure (Unseen)":"86.3","Jaccard (Seen)":"82.6","Jaccard (Unseen)":"78.3","O (Average of Measures)":"83.6"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2207.00887","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2207.00887"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/jerryx1110/robust-video-object-segmentation","reach":null}],"summary":{"ran_violates":1,"ran_fixture":1,"unverified":2},"by_repo_kind":{"official":{"samples":4,"ran":2,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":4,"samples":[{"code_sha256_prefix":"7221953c05ccc6b9","entry":"all_to_onehot","repo":"jerryx1110/robust-video-object-segmentation","repo_kind":"official","path":"Robust-VOS-Benchmark/CFBI&AOC(ours)/datasets_robustness.py","file_url":"https://github.com/jerryx1110/robust-video-object-segmentation/blob/HEAD/Robust-VOS-Benchmark/CFBI%26AOC%28ours%29/datasets_robustness.py","link_basis":"first_harvest_node","language":"python","status":"ran_violates","verification_level":1,"contract_check":"VIOLATES","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"7221953c05ccc6b9"}},{"code_sha256_prefix":"ba043a4f7b01fc06","entry":"foreground2background","repo":"jerryx1110/robust-video-object-segmentation","repo_kind":"official","path":"AOC-Net/adaptive_embedding_for_matching.py","file_url":"https://github.com/jerryx1110/robust-video-object-segmentation/blob/HEAD/AOC-Net/adaptive_embedding_for_matching.py","link_basis":"first_harvest_node","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"RAISES","metamorphic_tier":"well_formed","behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"ba043a4f7b01fc06"}},{"code_sha256_prefix":"1e1739fe8e406dc0","entry":"global_matching_cluster","repo":"jerryx1110/robust-video-object-segmentation","repo_kind":"official","path":"AOC-Net/adaptive_embedding_for_matching.py","file_url":"https://github.com/jerryx1110/robust-video-object-segmentation/blob/HEAD/AOC-Net/adaptive_embedding_for_matching.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"1e1739fe8e406dc0"}},{"code_sha256_prefix":"bdcb19540ed18157","entry":"global_matching_proxy","repo":"jerryx1110/robust-video-object-segmentation","repo_kind":"official","path":"AOC-Net/adaptive_embedding_for_matching.py","file_url":"https://github.com/jerryx1110/robust-video-object-segmentation/blob/HEAD/AOC-Net/adaptive_embedding_for_matching.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"bdcb19540ed18157"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}