Browse State-of-the-Art › Video Object Segmentation
Video Object Segmentation
294 papers with code · 13 benchmarks · 19 datasets archive 2025-07-28
Video object segmentation is a binary labeling problem aiming to separate foreground object(s) from the background region of a video.
For leaderboards please refer to the different subtasks.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
13 leaderboard tables shown for this task, 13 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 13 until expanded.
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
19 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
7 subtasks in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 294 papers with code (551 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
29 Apr 2021 32 repositories listed Syntology ran 5 of 20 samples · 15 unverified · 2 pointer-only (licence)In this paper, we question if self-supervised learning provides new properties to Vision Transformer (ViT) that stand out compared to convolutional networks (convnets).
-
1 Aug 2024 11 repositories listed Syntology ran 28 of 49 samples · 21 unverifiedWe present Segment Anything Model 2 (SAM 2), a foundation model towards solving promptable visual segmentation in images and videos.
-
16 Nov 2016 8 repositories listedThis paper tackles the task of semi-supervised video object segmentation, i.
-
31 Mar 2021 5 repositories listed Syntology ran 1 of 1 samples · 0 unverifiedTo learn generalizable representation for correspondence in large-scale, a variety of self-supervised pretext tasks are proposed to explicitly perform object-level or patch-level similarity learning.
-
14 Mar 2021 5 repositories listed Syntology ran 9 of 18 samples · 9 unverifiedWe present Modular interactive VOS (MiVOS) framework which decouples interaction-to-mask and mask propagation, allowing for higher generalizability and better performance.
-
24 Jul 2018 5 repositories listed Syntology ran 3 of 15 samples · 12 unverified · 3 pointer-only (licence)We address semi-supervised video object segmentation, the task of automatically generating accurate and consistent pixel masks for objects in a video sequence, given the first-frame ground truth annotations.
-
27 Mar 2022 4 repositories listed Syntology ran 0 of 8 samples · 8 unverifiedWe present the first comprehensive video polyp segmentation (VPS) study in the deep learning era.
-
3 Dec 2020 4 repositories listedIn the semi-supervised setting, the first mask of each object is provided at test time.
-
16 Jul 2020 4 repositories listed Syntology ran 0 of 15 samples · 15 unverifiedThe global transfer module conveys the segmentation information in an annotated frame to a target frame, while the local transfer module propagates the segmentation information in a temporally adjacent frame to the…
-
3 Sep 2018 4 repositories listedEnd-to-end sequential learning to explore spatial-temporal features for video segmentation is largely limited by the scale of available video segmentation datasets, i.
-
28 Mar 2017 4 repositories listedOur approach is suitable for both single and multiple object segmentation.
-
13 Apr 2023 3 repositories listedIn SEEM, we propose a novel decoding mechanism that enables diverse prompting for all types of segmentation tasks, aiming at a universal segmentation interface that behaves like large language models (LLMs).
-
6 Apr 2023 3 repositories listed Syntology ran 6 of 10 samples · 4 unverifiedWe unify various segmentation tasks into a generalist in-context learning framework that accommodates different kinds of segmentation data by transforming them into the same format of images.
-
26 Sep 2022 3 repositories listed Syntology ran 7 of 7 samples · 0 unverified · 7 pointer-only (licence)VISOR annotates videos from EPIC-KITCHENS, which comes with a new set of challenges not encountered in current video segmentation datasets.
-
Rethinking Space-Time Networks with Improved Memory Coverage for Efficient Video Object Segmentation9 Jun 2021 3 repositories listed Syntology ran 6 of 10 samples · 4 unverified · 2 pointer-only (licence)This paper presents a simple yet effective approach to modeling space-time correspondences in the context of video object segmentation.
-
30 Apr 2020 3 repositories listedWe describe our development and show the use of our solver in a video segmentation task and meta-learning for few-shot learning.
-
29 Sep 2019 3 repositories listedTo handle the nonrigid background like a sea, we also propose a robust fusion mechanism between motion and appearance-based features.
-
1 Apr 2019 3 repositories listed Syntology ran 1 of 1 samples · 0 unverifiedIn our framework, the past frames with object masks form an external memory, and the current frame as the query is segmented using the mask information in the memory.
-
25 Feb 2019 3 repositories listed Syntology ran 1 of 9 samples · 8 unverifiedMany of the recent successful methods for video object segmentation (VOS) are overly complicated, heavily rely on fine-tuning on the first frame, and/or are slow, and are hence of limited practical use.
-
12 Dec 2018 3 repositories listed Syntology ran 1 of 10 samples · 9 unverifiedIn this paper we illustrate how to perform both visual object tracking and semi-supervised video object segmentation, in real-time, with a single simple approach.
-
1 Aug 2017 3 repositories listedSpecifically, our Video Object Segmentation with Re-identification (VS-ReID) model includes a mask propagation module and a ReID module.
-
3 May 2025 2 repositories listedManual annotation of volumetric medical images, such as magnetic resonance imaging (MRI) and computed tomography (CT), is a labor-intensive and time-consuming process.
-
10 Dec 2024 2 repositories listed Syntology ran 2 of 6 samples · 4 unverifiedIn this work, we focus on semi-supervised learning for video action detection.
-
16 Jul 2024 2 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)In this paper, we introduce a new task, Reasoning Video Object Segmentation (ReasonVOS).
-
24 Jun 2024 2 repositories listedMoreover, we provide a new motion expression guided video segmentation dataset MeViS to study the natural language-guided video understanding in complex environments.
-
25 Dec 2023 2 repositories listed Syntology ran 9 of 9 samples · 0 unverifiedWe evaluate our unified models on various benchmarks.
-
29 Jul 2023 2 repositories listed Syntology ran 0 of 1 samples · 1 unverified · 1 pointer-only (licence)Despite advancements in user-guided video segmentation, extracting complex objects consistently for highly complex scenes is still a labor-intensive task, especially for production.
-
21 Jun 2023 2 repositories listedOnline unsupervised video object segmentation (UVOS) uses the previous frames as its input to automatically separate the primary object(s) from a streaming video without using any further manual annotation.
-
8 May 2023 2 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)Considering the challenges in panoptic VOS, we propose a strong baseline method named panoptic object association with transformers (PAOT), which uses panoptic identification to associate objects with a pyramid…
-
28 Apr 2023 2 repositories listedSubsequently, we devise EmoFormer, a novel network able to exploit the event data.
Syntology lines on 18 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections