{"url":"/sota/semi-supervised-video-object-segmentation-on-20","task":{"name":"Semi-Supervised Video Object Segmentation","url":"/task/semi-supervised-video-object-segmentation","note":null},"dataset":{"name":"DAVIS (no YouTube-VOS training)","url":"/dataset/davis"},"category":"Computer Vision","categories":["Computer Vision"],"category_note":null,"description":"The semi-supervised scenario assumes the user inputs a full mask of the object(s) of interest in the first frame of a video sequence. Methods have to produce the segmentation mask for that object(s) in the subsequent frames.","description_from":"task","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","rank":"the archive's row order at snapshot; not re-ranked","rows_end_at":"2025-07-28","rows_withheld_as_spam":0,"metric_values":"the archive's strings, untouched"},"metrics":["D17 val (G)","D17 val (J)","D17 val (F)","D17 test (G)","D17 test (J)","D17 test (F)","D16 val (G)","D16 val (J)","D16 val (F)","FPS"],"metric_direction":{"note":"inferred from the metric name only (the archive records no direction); null = not inferred, chart draws points only","by_metric":{"D17 val (G)":null,"D17 val (J)":null,"D17 val (F)":null,"D17 test (G)":null,"D17 test (J)":null,"D17 test (F)":null,"D16 val (G)":null,"D16 val (J)":null,"D16 val (F)":null,"FPS":null}},"counts":{"rows":26,"rows_with_code":22,"rows_with_paper_page":25,"rows_dated":25,"rows_using_additional_data":0},"rows":[{"rank_in_archive_order":1,"model":"HMMN","metrics":{"D16 val (F)":"90.6","D16 val (G)":"89.4","D16 val (J)":"88.2","D17 val (F)":"83.1","D17 val (G)":"80.4","D17 val (J)":"77.7","FPS":"10.0"},"uses_additional_data":false,"paper_date":"2021-09-23","paper":"/paper/hierarchical-memory-matching-network-for","paper_url":"https://arxiv.org/abs/2109.11404v1","paper_title":"Hierarchical Memory Matching Network for Video Object Segmentation","code":"https://github.com/hongje/hmmn","n_code_links":1,"syntology":{"n_ran":5,"n_unverified":3,"n_samples":8,"n_pointer_only_licence":8}},{"rank_in_archive_order":2,"model":"TBD","metrics":{"D16 val (F)":"86.2","D16 val (G)":"86.8","D16 val (J)":"87.5","D17 test (F)":"72.2","D17 test (G)":"69.4","D17 test (J)":"66.6","D17 val (F)":"82.3","D17 val (G)":"80.0","D17 val (J)":"77.6","FPS":"50.1"},"uses_additional_data":false,"paper_date":"2022-07-14","paper":"/paper/tackling-background-distraction-in-video","paper_url":"https://arxiv.org/abs/2207.06953v3","paper_title":"Tackling Background Distraction in Video Object Segmentation","code":"https://github.com/suhwan-cho/tbd","n_code_links":1,"syntology":{"n_ran":1,"n_unverified":2,"n_samples":3,"n_pointer_only_licence":0}},{"rank_in_archive_order":3,"model":"AOT-S","metrics":{"D17 val (F)":"82.0","D17 val (G)":"79.2","D17 val (J)":"76.4","FPS":"40.0"},"uses_additional_data":false,"paper_date":"2021-06-04","paper":"/paper/associating-objects-with-transformers-for","paper_url":"https://arxiv.org/abs/2106.02638v3","paper_title":"Associating Objects with Transformers for Video Object Segmentation","code":"https://github.com/yoxu515/aot-benchmark","n_code_links":2,"syntology":null},{"rank_in_archive_order":4,"model":"JOINT","metrics":{"D17 val (F)":"81.2","D17 val (G)":"78.6","D17 val (J)":"76.0","FPS":"4.00"},"uses_additional_data":false,"paper_date":"2021-08-08","paper":"/paper/joint-inductive-and-transductive-learning-for","paper_url":"https://arxiv.org/abs/2108.03679v1","paper_title":"Joint Inductive and Transductive Learning for Video Object Segmentation","code":"https://github.com/maoyunyao/joint","n_code_links":1,"syntology":{"n_ran":0,"n_unverified":1,"n_samples":1,"n_pointer_only_licence":1}},{"rank_in_archive_order":5,"model":"SSTVOS","metrics":{"D17 val (F)":"81.4","D17 val (G)":"78.4","D17 val (J)":"75.4"},"uses_additional_data":false,"paper_date":"2021-01-21","paper":"/paper/sstvos-sparse-spatiotemporal-transformers-for","paper_url":"https://arxiv.org/abs/2101.08833v2","paper_title":"SSTVOS: Sparse Spatiotemporal Transformers for Video Object Segmentation","code":"https://github.com/dukebw/SSTVOS","n_code_links":1,"syntology":{"n_ran":5,"n_unverified":2,"n_samples":7,"n_pointer_only_licence":7}},{"rank_in_archive_order":6,"model":"SWEM","metrics":{"D16 val (F)":"89.0","D16 val (G)":"88.1","D16 val (J)":"87.3","D17 val (F)":"79.8","D17 val (G)":"77.2","D17 val (J)":"74.5","FPS":"36.0"},"uses_additional_data":false,"paper_date":"2022-08-22","paper":"/paper/swem-towards-real-time-video-object-1","paper_url":"https://arxiv.org/abs/2208.10128v1","paper_title":"SWEM: Towards Real-Time Video Object Segmentation with Sequential Weighted Expectation-Maximization","code":"https://github.com/lmm077/SWEM","n_code_links":1,"syntology":null},{"rank_in_archive_order":7,"model":"KMN","metrics":{"D16 val (F)":"88.1","D16 val (G)":"87.6","D16 val (J)":"87.1","D17 val (F)":"77.8","D17 val (G)":"76.0","D17 val (J)":"74.2","FPS":"8.33"},"uses_additional_data":false,"paper_date":"2020-07-16","paper":"/paper/kernelized-memory-network-for-video-object","paper_url":"https://arxiv.org/abs/2007.08270v1","paper_title":"Kernelized Memory Network for Video Object Segmentation","code":"https://github.com/hkchengrex/Mask-Propagation","n_code_links":1,"syntology":null},{"rank_in_archive_order":8,"model":"LCM","metrics":{"D17 val (F)":"77.2","D17 val (G)":"75.2","D17 val (J)":"73.1","FPS":"8.47"},"uses_additional_data":false,"paper_date":"2021-04-09","paper":"/paper/learning-position-and-target-consistency-for","paper_url":"https://arxiv.org/abs/2104.04329v1","paper_title":"Learning Position and Target Consistency for Memory-based Video Object Segmentation","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":9,"model":"RMNet","metrics":{"D16 val (F)":"82.3","D16 val (G)":"81.5","D16 val (J)":"80.6","D17 val (F)":"77.2","D17 val (G)":"75.0","D17 val (J)":"72.8","FPS":"11.9"},"uses_additional_data":false,"paper_date":"2021-03-24","paper":"/paper/efficient-regional-memory-network-for-video","paper_url":"https://arxiv.org/abs/2103.12934v2","paper_title":"Efficient Regional Memory Network for Video Object Segmentation","code":"https://github.com/hzxie/RMNet","n_code_links":1,"syntology":{"n_ran":5,"n_unverified":5,"n_samples":10,"n_pointer_only_licence":0}},{"rank_in_archive_order":10,"model":"CFBI","metrics":{"D16 val (F)":"86.9","D16 val (G)":"86.1","D16 val (J)":"85.3","D17 val (F)":"77.7","D17 val (G)":"74.9","D17 val (J)":"72.1","FPS":"5.56"},"uses_additional_data":false,"paper_date":"2020-03-18","paper":"/paper/collaborative-video-object-segmentation-by","paper_url":"https://arxiv.org/abs/2003.08333v2","paper_title":"Collaborative Video Object Segmentation by Foreground-Background Integration","code":"https://github.com/PaddlePaddle/PaddleVideo/blob/develop/docs/en/model_zoo/segmentation/cfbi.md","n_code_links":2,"syntology":null},{"rank_in_archive_order":11,"model":"STG-Net","metrics":{"D16 val (F)":"86.0","D16 val (G)":"85.7","D16 val (J)":"85.4","D17 test (F)":"66.5","D17 test (G)":"63.1","D17 test (J)":"59.7","D17 val (F)":"77.9","D17 val (G)":"74.7","D17 val (J)":"71.5"},"uses_additional_data":false,"paper_date":"2020-12-10","paper":"/paper/spatiotemporal-graph-neural-network-based","paper_url":"https://arxiv.org/abs/2012.05499v1","paper_title":"Spatiotemporal Graph Neural Network based Mask Reconstruction for Video Object Segmentation","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":12,"model":"AFB-URR","metrics":{"D17 val (F)":"76.1","D17 val (G)":"74.6","D17 val (J)":"73.0","FPS":"4.00"},"uses_additional_data":false,"paper_date":"2020-10-15","paper":"/paper/video-object-segmentation-with-adaptive","paper_url":"https://arxiv.org/abs/2010.07958v1","paper_title":"Video Object Segmentation with Adaptive Feature Bank and Uncertain-Region Refinement","code":"https://github.com/xmlyqing00/AFB-URR","n_code_links":1,"syntology":null},{"rank_in_archive_order":13,"model":"LWL","metrics":{"D17 val (F)":"76.3","D17 val (G)":"74.3","D17 val (J)":"72.2","FPS":"14.0"},"uses_additional_data":false,"paper_date":"2020-03-25","paper":"/paper/learning-what-to-learn-for-video-object","paper_url":"https://arxiv.org/abs/2003.11540v2","paper_title":"Learning What to Learn for Video Object Segmentation","code":"https://github.com/visionml/pytracking","n_code_links":2,"syntology":null},{"rank_in_archive_order":14,"model":"BMVOS","metrics":{"D16 val (F)":"81.4","D16 val (G)":"82.2","D16 val (J)":"82.9","D17 test (F)":"64.7","D17 test (G)":"62.7","D17 test (J)":"60.7","D17 val (F)":"74.7","D17 val (G)":"72.7","D17 val (J)":"70.7","FPS":"45.9"},"uses_additional_data":false,"paper_date":"2021-10-04","paper":"/paper/pixel-level-bijective-matching-for-video","paper_url":"https://arxiv.org/abs/2110.01644v3","paper_title":"Pixel-Level Bijective Matching for Video Object Segmentation","code":"https://github.com/suhwan-cho/bmvos","n_code_links":1,"syntology":null},{"rank_in_archive_order":15,"model":"TVOS","metrics":{"D17 test (F)":"67.4","D17 test (G)":"63.1","D17 test (J)":"58.8","D17 val (F)":"74.7","D17 val (G)":"72.3","D17 val (J)":"69.9","FPS":"37.0"},"uses_additional_data":false,"paper_date":"2020-04-15","paper":"/paper/a-transductive-approach-for-video-object","paper_url":"https://arxiv.org/abs/2004.07193v2","paper_title":"A Transductive Approach for Video Object Segmentation","code":"https://github.com/microsoft/transductive-vos.pytorch","n_code_links":1,"syntology":null},{"rank_in_archive_order":16,"model":"STM","metrics":{"D16 val (F)":"88.1","D16 val (G)":"86.5","D16 val (J)":"84.8","D17 val (F)":"74.0","D17 val (G)":"71.6","D17 val (J)":"69.2","FPS":"6.25"},"uses_additional_data":false,"paper_date":"2019-04-01","paper":"/paper/video-object-segmentation-using-space-time","paper_url":"https://arxiv.org/abs/1904.00607v2","paper_title":"Video Object Segmentation using Space-Time Memory Networks","code":"https://github.com/seoungwugoh/STM","n_code_links":3,"syntology":{"n_ran":1,"n_unverified":0,"n_samples":1,"n_pointer_only_licence":0}},{"rank_in_archive_order":17,"model":"GC","metrics":{"D16 val (F)":"85.7","D16 val (G)":"86.6","D16 val (J)":"87.6","D17 val (F)":"73.5","D17 val (G)":"71.4","D17 val (J)":"69.3","FPS":"25.0"},"uses_additional_data":false,"paper_date":"2020-01-30","paper":"/paper/fast-video-object-segmentation-using-the","paper_url":"https://arxiv.org/abs/2001.11243v2","paper_title":"Fast Video Object Segmentation using the Global Context Module","code":"https://github.com/cmsflash/global-context-module","n_code_links":1,"syntology":null},{"rank_in_archive_order":18,"model":"DMM-Net","metrics":{"D17 val (F)":"73.3","D17 val (G)":"70.7","D17 val (J)":"68.1"},"uses_additional_data":false,"paper_date":"2019-09-27","paper":"/paper/dmm-net-differentiable-mask-matching-network","paper_url":"https://arxiv.org/abs/1909.12471v1","paper_title":"DMM-Net: Differentiable Mask-Matching Network for Video Object Segmentation","code":"https://github.com/ZENGXH/DMM_Net","n_code_links":1,"syntology":null},{"rank_in_archive_order":19,"model":"FEELVOS","metrics":{"D16 val (F)":"83.1","D16 val (G)":"81.7","D16 val (J)":"80.3","D17 test (F)":"57.5","D17 test (G)":"54.4","D17 test (J)":"51.2","D17 val (F)":"72.3","D17 val (G)":"69.1","D17 val (J)":"65.9","FPS":"2.22"},"uses_additional_data":false,"paper_date":"2019-02-25","paper":"/paper/feelvos-fast-end-to-end-embedding-learning","paper_url":"http://arxiv.org/abs/1902.09513v2","paper_title":"FEELVOS: Fast End-to-End Embedding Learning for Video Object Segmentation","code":"https://github.com/tensorflow/models","n_code_links":3,"syntology":{"n_ran":9,"n_unverified":0,"n_samples":9,"n_pointer_only_licence":0}},{"rank_in_archive_order":20,"model":"FRTM","metrics":{"D16 val (G)":"81.7","D17 val (F)":"71.2","D17 val (G)":"68.8","D17 val (J)":"66.4","FPS":"21.9"},"uses_additional_data":false,"paper_date":"2020-02-27","paper":"/paper/learning-fast-and-robust-target-models-for","paper_url":"https://arxiv.org/abs/2003.00908v2","paper_title":"Learning Fast and Robust Target Models for Video Object Segmentation","code":"https://github.com/andr345/frtm-vos","n_code_links":2,"syntology":null},{"rank_in_archive_order":21,"model":"DIPNet","metrics":{"D16 val (F)":"86.4","D16 val (G)":"86.1","D16 val (J)":"85.8","D17 test (G)":"55.2","D17 val (F)":"71.6","D17 val (G)":"68.5","D17 val (J)":"65.3","FPS":"0.92"},"uses_additional_data":false,"paper_date":null,"paper":null,"paper_url":null,"paper_title":"","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":22,"model":"AGSS-VOS","metrics":{"D17 test (F)":"59.7","D17 test (G)":"57.2","D17 test (J)":"54.8","D17 val (F)":"69.9","D17 val (G)":"67.4","D17 val (J)":"64.9","FPS":"10.0"},"uses_additional_data":false,"paper_date":"2019-10-01","paper":"/paper/agss-vos-attention-guided-single-shot-video","paper_url":"http://openaccess.thecvf.com/content_ICCV_2019/html/Lin_AGSS-VOS_Attention_Guided_Single-Shot_Video_Object_Segmentation_ICCV_2019_paper.html","paper_title":"AGSS-VOS: Attention Guided Single-Shot Video Object Segmentation","code":"https://github.com/Jia-Research-Lab/AGSS-VOS","n_code_links":1,"syntology":null},{"rank_in_archive_order":23,"model":"DTN","metrics":{"D16 val (F)":"83.5","D16 val (G)":"83.6","D16 val (J)":"83.7","D17 val (F)":"70.6","D17 val (G)":"67.4","D17 val (J)":"64.2","FPS":"14.3"},"uses_additional_data":false,"paper_date":"2019-10-01","paper":"/paper/fast-video-object-segmentation-via-dynamic","paper_url":"http://openaccess.thecvf.com/content_ICCV_2019/html/Zhang_Fast_Video_Object_Segmentation_via_Dynamic_Targeting_Network_ICCV_2019_paper.html","paper_title":"Fast Video Object Segmentation via Dynamic Targeting Network","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":24,"model":"RANet","metrics":{"D16 val (F)":"85.4","D16 val (G)":"85.5","D16 val (J)":"85.5","D17 test (F)":"57.2","D17 test (G)":"55.3","D17 test (J)":"53.4","D17 val (F)":"68.2","D17 val (G)":"65.7","D17 val (J)":"63.2","FPS":"30.3"},"uses_additional_data":false,"paper_date":"2019-08-19","paper":"/paper/ranet-ranking-attention-network-for-fast","paper_url":"https://arxiv.org/abs/1908.06647v4","paper_title":"RANet: Ranking Attention Network for Fast Video Object Segmentation","code":"https://github.com/Storife/RANet","n_code_links":2,"syntology":{"n_ran":4,"n_unverified":8,"n_samples":12,"n_pointer_only_licence":0}},{"rank_in_archive_order":25,"model":"STCNN","metrics":{"D16 val (F)":"83.8","D16 val (G)":"83.8","D16 val (J)":"83.8","D17 val (F)":"64.6","D17 val (G)":"61.7","D17 val (J)":"58.7","FPS":"0.26"},"uses_additional_data":false,"paper_date":"2019-04-04","paper":"/paper/spatiotemporal-cnn-for-video-object","paper_url":"http://arxiv.org/abs/1904.02363v1","paper_title":"Spatiotemporal CNN for Video Object Segmentation","code":"https://github.com/longyin880815/STCNN","n_code_links":1,"syntology":{"n_ran":1,"n_unverified":0,"n_samples":1,"n_pointer_only_licence":1}},{"rank_in_archive_order":26,"model":"XMem","metrics":{"FPS":"29.6"},"uses_additional_data":false,"paper_date":"2022-07-14","paper":"/paper/xmem-long-term-video-object-segmentation-with","paper_url":"https://arxiv.org/abs/2207.07115v2","paper_title":"XMem: Long-Term Video Object Segmentation with an Atkinson-Shiffrin Memory Model","code":"https://github.com/hkchengrex/XMem","n_code_links":2,"syntology":{"n_ran":1,"n_unverified":2,"n_samples":3,"n_pointer_only_licence":2}}],"since_archive":{"claim":"Results that newer papers report for their own method, placed here by Syntology. A model pointed at the cell in the paper's own table; the number was read from that cell and checked against this leaderboard's metric, dataset, split and scale; an independent check that saw this leaderboard's other rows and every other leaderboard on the same dataset accepted it. Not reviewed by the paper's authors or by the archive's editors, and not ranked against the archive rows.","extraction_file_present":true,"measurement":{"test_papers":883,"papers_with_output":881,"judged_true":108,"judged":110,"wilson95_lower":0.9361,"measured_on":"2026-09-24","frozen_commit":"0e3de0df94"},"measurement_note":"blind adjudication of accepted entries on a held-out split of archive papers, rules frozen before the test","coverage":{"sentence":"Syntology has checked 6,885 of the 9,623 papers on this site that are newer than the archive; results from the others appear after they are checked.","complete":false,"papers_newer_than_archive":9623,"papers_checked":6885,"papers_extracted_not_yet_verified":0,"boards_without_verdict":2,"papers_not_yet_extracted":2737},"order":"newest first by month (arXiv date, else the arXiv-id month), then arXiv id descending","columns":[],"entries":[]},"syntology":{"read_at":"2026-09-25T09:33:49+00:00","claim":"Per row: N of M harvested code samples from that row's paper executed on a synthesized fixture; the other M-N are unverified. Not a reproduction of the row's number; not a correctness claim. n_pointer_only_licence counts samples the site points at rather than redistributes (a licence axis, independent of ran/unverified).","rows_with_graph_line":10,"rows_with_any_sample_ran":9,"distinct_papers_with_graph_line":10,"distinct_papers_with_any_sample_ran":9,"samples_over_distinct_papers":{"n_ran":32,"n_unverified":23,"n_samples":55,"n_pointer_only_licence":19,"note":"each paper (arXiv id) counted once, however many rows it is behind; this is the page-level figure"},"samples_row_weighted":{"n_ran":32,"n_unverified":23,"n_samples":55,"n_pointer_only_licence":19,"note":"row-weighted: a paper behind several rows is counted once per row; inflated relative to samples_over_distinct_papers by design, kept for readers summing the per-row syntology blocks"}}}