Browse State-of-the-Art › Semi-Supervised Video Object Segmentation
Semi-Supervised Video Object Segmentation
99 papers with code · 16 benchmarks · 13 datasets archive 2025-07-28
The semi-supervised scenario assumes the user inputs a full mask of the object(s) of interest in the first frame of a video sequence. Methods have to produce the segmentation mask for that object(s) in the subsequent frames.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
16 leaderboard tables shown for this task, 16 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 16 until expanded.
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
13 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
1 subtask in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 99 papers with code (147 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
1 Aug 2024 11 repositories listed Syntology ran 28 of 49 samples · 21 unverifiedWe present Segment Anything Model 2 (SAM 2), a foundation model towards solving promptable visual segmentation in images and videos.
-
16 Nov 2016 8 repositories listedThis paper tackles the task of semi-supervised video object segmentation, i.
-
14 Mar 2021 5 repositories listed Syntology ran 9 of 18 samples · 9 unverifiedWe present Modular interactive VOS (MiVOS) framework which decouples interaction-to-mask and mask propagation, allowing for higher generalizability and better performance.
-
24 Jul 2018 5 repositories listed Syntology ran 3 of 15 samples · 12 unverified · 3 pointer-only (licence)We address semi-supervised video object segmentation, the task of automatically generating accurate and consistent pixel masks for objects in a video sequence, given the first-frame ground truth annotations.
-
3 Dec 2020 4 repositories listedIn the semi-supervised setting, the first mask of each object is provided at test time.
-
3 Sep 2018 4 repositories listedEnd-to-end sequential learning to explore spatial-temporal features for video segmentation is largely limited by the scale of available video segmentation datasets, i.
-
28 Mar 2017 4 repositories listedOur approach is suitable for both single and multiple object segmentation.
-
Rethinking Space-Time Networks with Improved Memory Coverage for Efficient Video Object Segmentation9 Jun 2021 3 repositories listed Syntology ran 6 of 10 samples · 4 unverified · 2 pointer-only (licence)This paper presents a simple yet effective approach to modeling space-time correspondences in the context of video object segmentation.
-
1 Apr 2019 3 repositories listed Syntology ran 1 of 1 samples · 0 unverifiedIn our framework, the past frames with object masks form an external memory, and the current frame as the query is segmented using the mask information in the memory.
-
25 Feb 2019 3 repositories listed Syntology ran 1 of 9 samples · 8 unverifiedMany of the recent successful methods for video object segmentation (VOS) are overly complicated, heavily rely on fine-tuning on the first frame, and/or are slow, and are hence of limited practical use.
-
12 Dec 2018 3 repositories listed Syntology ran 1 of 10 samples · 9 unverifiedIn this paper we illustrate how to perform both visual object tracking and semi-supervised video object segmentation, in real-time, with a single simple approach.
-
29 Jul 2023 2 repositories listed Syntology ran 0 of 1 samples · 1 unverified · 1 pointer-only (licence)Despite advancements in user-guided video segmentation, extracting complex objects consistently for highly complex scenes is still a labor-intensive task, especially for production.
-
8 May 2023 2 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)Considering the challenges in panoptic VOS, we propose a strong baseline method named panoptic object association with transformers (PAOT), which uses panoptic identification to associate objects with a pyramid…
-
18 Oct 2022 2 repositories listedTo solve such a problem and further facilitate the learning of visual embeddings, this paper proposes a Decoupling Features in Hierarchical Propagation (DeAOT) approach.
-
14 Jul 2022 2 repositories listed Syntology ran 1 of 3 samples · 2 unverified · 2 pointer-only (licence)We present XMem, a video object segmentation architecture for long videos with unified feature memory stores inspired by the Atkinson-Shiffrin memory model.
-
22 Mar 2022 2 repositories listedThis paper delves into the challenges of achieving scalable and effective multi-object modeling for semi-supervised Video Object Segmentation (VOS).
-
26 Jul 2021 2 repositories listed Syntology ran 0 of 2 samples · 2 unverified · 2 pointer-only (licence)We propose an efficient plug-and-play acceleration framework for semi-supervised video object segmentation by exploiting the temporal redundancies in videos presented by the compressed bitstream.
-
4 Jun 2021 2 repositories listedThe state-of-the-art methods learn to decode features with a single positive object and thus have to match and segment each target separately under multi-object scenarios, consuming multiple times computing resources.
-
25 Mar 2020 2 repositories listedThis allows us to achieve a rich internal representation of the target in the current frame, significantly increasing the segmentation accuracy of our approach.
-
18 Mar 2020 2 repositories listedThis paper investigates the principles of embedding learning to tackle the challenging semi-supervised video object segmentation.
-
27 Feb 2020 2 repositories listedThe target appearance model consists of a light-weight module, which is learned during the inference stage using fast optimization techniques to predict a coarse but robust target segmentation.
-
18 Feb 2020 2 repositories listedRecent interest in self-supervised dense tracking has yielded rapid progress, but performance still remains far from supervised methods.
-
26 Sep 2019 2 repositories listed Syntology ran 4 of 23 samples · 19 unverifiedOur learning process integrates two highly related tasks: tracking large image regions \emph{and} establishing fine-grained pixel-level associations between consecutive video frames.
-
19 Aug 2019 2 repositories listed Syntology ran 1 of 12 samples · 11 unverifiedSpecifically, to integrate the insights of matching based and propagation based methods, we employ an encoder-decoder framework to learn pixel-level similarity and segmentation in an end-to-end manner.
-
6 Jun 2018 2 repositories listed Syntology ran 0 of 3 samples · 3 unverified · 3 pointer-only (licence)基于视频的目标检测算法研究
-
1 Jun 2018 2 repositories listedWe validate our method on four benchmark sets that cover single and multiple object segmentation.
-
8 Dec 2016 2 repositories listed Syntology ran 3 of 8 samples · 5 unverifiedInspired by recent advances of deep learning in instance segmentation and object tracking, we introduce video object segmentation problem as a concept of guided instance segmentation.
-
15 Dec 2024 1 repository listedContextual information at the video level has become increasingly crucial for visual object tracking.
-
26 Nov 2024 1 repository listed Syntology ran 3 of 11 samples · 8 unverified · 11 pointer-only (licence)We argue that a more sophisticated memory model is required, and propose a new distractor-aware memory model for SAM2 and an introspection-based update strategy that jointly addresses the segmentation accuracy as well…
-
5 Nov 2024 1 repository listedSemi-supervised video object segmentation (VOS) has been largely driven by space-time memory (STM) networks, which store past frame features in a spatiotemporal memory to segment the current frame via softmax attention.
Syntology lines on 16 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections