Browse State-of-the-Art › Unsupervised Video Object Segmentation
Unsupervised Video Object Segmentation
52 papers with code · 6 benchmarks · 8 datasets archive 2025-07-28
The unsupervised scenario assumes that the user does not interact with the algorithm to obtain the segmentation masks. Methods should provide a set of object candidates with no overlapping pixels that span through the whole video sequence. This set of objects should contain at least the objects that capture human attention when watching the whole video sequence i.e objects that are more likely to be followed by human gaze.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
6 leaderboard tables shown for this task, 6 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| DAVIS 2016 val (25 rows) | GSANet | Guided Slot Attention for Unsupervised Video Object Segmentation | code | — | Compare |
| YouTube-Objects (16 rows) | FakeFlow | Improving Unsupervised Video Object Segmentation via Fake Flow Generation | — | — | Compare |
| FBMS test (15 rows) | FakeFlow | Improving Unsupervised Video Object Segmentation via Fake Flow Generation | — | — | Compare |
| DAVIS 2017 (val) (10 rows) | DEVA (EntitySeg) | Tracking Anything with Decoupled Video Segmentation | code | Syntology ran 7 of 10 samples · 3 unverified | Compare |
| DAVIS 2017 (test-dev) (6 rows) | DEVA (EntitySeg) | Tracking Anything with Decoupled Video Segmentation | code | Syntology ran 7 of 10 samples · 3 unverified | Compare |
| SegTrack v2 (4 rows) | FrameSelect | Mask Selection and Propagation for Unsupervised Video Object Segmentation | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
8 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 52 papers with code (89 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
29 Sep 2019 3 repositories listedTo handle the nonrigid background like a sea, we also propose a robust fusion mechanism between motion and appearance-based features.
-
21 Jun 2023 2 repositories listedOnline unsupervised video object segmentation (UVOS) uses the previous frames as its input to automatically separate the primary object(s) from a streaming video without using any further manual annotation.
-
4 Sep 2022 2 repositories listed Syntology ran 0 of 4 samples · 4 unverifiedUnsupervised video object segmentation (VOS) aims to detect the most salient object in a video sequence at the pixel level.
-
15 Nov 2021 2 repositories listedWe further show that D2Conv3D out-performs trivial extensions of existing dilated and deformable convolutions to 3D.
-
18 Feb 2020 2 repositories listedRecent interest in self-supervised dense tracking has yielded rapid progress, but performance still remains far from supervised methods.
-
26 Sep 2019 2 repositories listed Syntology ran 4 of 23 samples · 19 unverifiedOur learning process integrates two highly related tasks: tracking large image regions \emph{and} establishing fine-grained pixel-level associations between consecutive video frames.
-
19 Nov 2018 2 repositories listedWe investigate the problem of strictly unsupervised video object segmentation, i.
-
14 Jan 2025 1 repository listedIn this paper, we address the challenges in unsupervised video object segmentation (UVOS) by proposing an efficient algorithm, termed MTNet, which concurrently exploits motion and temporal cues.
-
26 Sep 2023 1 repository listedUnsupervised video object segmentation (VOS) is a task that aims to detect the most salient object in a video without external guidance about the object.
-
7 Sep 2023 1 repository listed Syntology ran 7 of 10 samples · 3 unverified · 10 pointer-only (licence)To 'track anything' without training on video data for every individual task, we develop a decoupled video segmentation approach (DEVA), composed of task-specific image-level segmentation and class/task-agnostic…
-
22 May 2023 1 repository listedExtensive experiments on the DAVIS2017-unsupervised and YoutubeVIS19\&21 datasets demonstrate the superior performance of UVOSAM without mask supervision compared to existing mask-supervised methods, as well as its…
-
17 Apr 2023 1 repository listedThe Gestalt law of common fate, i.
-
18 Mar 2023 1 repository listedIn the static object predictor, the RGB source is converted to depth and static saliency sources, simultaneously.
-
15 Mar 2023 1 repository listedUnsupervised video object segmentation aims to segment the most prominent object in a video sequence.
-
22 Nov 2022 1 repository listedUnsupervised video object segmentation (VOS) aims to detect and segment the most salient object in videos.
-
19 Sep 2022 1 repository listed Syntology ran 4 of 10 samples · 6 unverifiedWe propose a simple, yet powerful approach for unsupervised object segmentation in videos.
-
8 Sep 2022 1 repository listedThe proposed model effectively extracts the RGB and motion information by extracting superpixel-based component prototypes from the input RGB images and optical flow maps.
-
18 Jul 2022 1 repository listedOptical flow is an easily conceived and precious cue for advancing unsupervised video object segmentation (UVOS).
-
6 Apr 2022 1 repository listedUnsupervised video object segmentation (UVOS) aims at automatically separating the primary foreground object(s) from the background in a video sequence.
-
29 Mar 2022 1 repository listedBy contrast, pixel-level optimization is more explicit, however, it is sensitive to the visual quality of training data and is not robust to object deformation.
-
15 Dec 2021 1 repository listedThe main novelty of the proposed model is that the autoencoder is also trained to predict the background noise, which allows to compute for each frame a pixel-dependent threshold to perform the foreground segmentation.
-
15 Nov 2021 1 repository listedWe further show that D^2Conv3D out-performs trivial extensions of existing dilated and deformable convolutions to 3D.
-
11 Nov 2021 1 repository listed Syntology ran 6 of 6 samples · 0 unverifiedOn established VOS benchmarks, our approach exceeds the segmentation accuracy of previous work despite using significantly less training data and compute power.
-
11 Aug 2021 1 repository listedIn this paper, we propose a novel multi-source fusion network for zero-shot video object segmentation.
-
6 Aug 2021 1 repository listed Syntology ran 7 of 8 samples · 1 unverifiedPrevious video object segmentation approaches mainly focus on using simplex solutions between appearance and motion, limiting feature collaboration efficiency among and across these two cues.
-
19 Jun 2021 1 repository listedAdditionally, to exclude the information of the moving background objects from motion features, our transformation module enables to reciprocally transform the appearance features to enhance the motion features, so as…
-
25 Mar 2021 1 repository listed Syntology ran 2 of 5 samples · 3 unverified · 5 pointer-only (licence)Video instance segmentation (VIS) aims to segment and associate all instances of predefined classes for each frame in videos.
-
5 Jan 2021 1 repository listedWe efficiently handle problems present in existing methods such as drift while temporal propagation, tracking and addition of new objects.
-
1 Jan 2021 1 repository listedHow to make the appearance and motion information interact effectively to accommodate complex scenarios is a fundamental issue in flow-based zero-shot video object segmentation.
-
26 Aug 2020 1 repository listed Syntology ran 3 of 9 samples · 6 unverifiedOn the other hand, 3D convolutional networks have been successfully applied for video classification tasks, but have not been leveraged as effectively to problems involving dense per-pixel interpretation of videos…
Syntology lines on 8 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections