Browse State-of-the-Art › Video Instance Segmentation
Video Instance Segmentation
94 papers with code · 8 benchmarks · 8 datasets archive 2025-07-28
The goal of video instance segmentation is simultaneous detection, segmentation and tracking of instances in videos. In words, it is the first time that the image instance segmentation problem is extended to the video domain.
To facilitate research on this new task, a large-scale benchmark called YouTube-VIS, which consists of 2,883 high-resolution YouTube videos, a 40-category label set and 131k high-quality instance masks is built.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
8 leaderboard tables shown for this task, 8 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
8 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 94 papers with code (148 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
21 Mar 2017 75 repositories listed Syntology ran 12 of 40 samples · 28 unverified · 7 pointer-only (licence)Simple Online and Realtime Tracking (SORT) is a pragmatic approach to multiple object tracking with a focus on simple, effective algorithms.
-
20 Dec 2021 6 repositories listed Syntology ran 2 of 7 samples · 5 unverifiedWe find Mask2Former also achieves state-of-the-art performance on video instance segmentation without modifying the architecture, the loss or even the training pipeline.
-
12 May 2019 6 repositories listed Syntology ran 3 of 11 samples · 8 unverified · 1 pointer-only (licence)The goal of this new task is simultaneous detection, segmentation and tracking of instances in videos.
-
5 May 2021 5 repositories listed Syntology ran 1 of 3 samples · 2 unverifiedThe key insight of QueryInst is to leverage the intrinsic one-to-one correspondence in object queries across different stages, as well as one-to-one correspondence between mask RoI features and object queries in the…
-
29 Mar 2024 3 repositories listed Syntology ran 1 of 5 samples · 4 unverified · 3 pointer-only (licence)Modern video segmentation methods adopt object queries to perform inter-frame association and demonstrate satisfactory performance in tracking continuously appearing objects despite large-scale motion and transient…
-
18 Apr 2022 3 repositories listed Syntology ran 0 of 2 samples · 2 unverifiedTo effectively and efficiently model the crucial temporal information within a video clip, we propose a Temporally Efficient Vision Transformer (TeViT) for video instance segmentation (VIS).
-
22 Mar 2024 2 repositories listedWe introduce InternVideo2, a new family of video foundation models (ViFM) that achieve the state-of-the-art results in video recognition, video-text tasks, and video-centric dialogue.
-
4 Feb 2024 2 repositories listedThen, these video prompts are prepended to the patch embeddings of the current frame as the updated input for video feature extraction.
-
3 Aug 2022 2 repositories listedBy only training a query-based image instance segmentation model, MinVIS outperforms the previous best result on the challenging Occluded VIS dataset by over 10% AP.
-
21 Jul 2022 2 repositories listed Syntology ran 3 of 5 samples · 2 unverifiedIn recent years, video instance segmentation (VIS) has been largely advanced by offline models, while online models gradually attracted less attention possibly due to their inferior performance.
-
8 Mar 2022 2 repositories listed Syntology ran 2 of 11 samples · 9 unverifiedGiven an input image or video, our framework first conducts multi-label classification over the complete label, then sorts the complete label and selects a small subset according to their class confidence scores.
-
15 Dec 2021 2 repositories listedNevertheless, we observe that a stand-alone instance query suffices for capturing a time sequence of instances in a video, but attention mechanisms shall be done with each frame independently.
-
15 Nov 2021 2 repositories listedWe further show that D2Conv3D out-performs trivial extensions of existing dilated and deformable convolutions to 3D.
-
22 Oct 2021 2 repositories listedIn this report, we introduce our (pretty straightforard) two-step "detect-then-match" video instance segmentation method.
-
10 Jun 2021 2 repositories listedContrastive self-supervised learning has outperformed supervised pretraining on many downstream tasks like segmentation and object detection.
-
2 Feb 2021 2 repositories listedOn the OVIS dataset, the highest AP achieved by state-of-the-art algorithms is only 16.
-
30 Nov 2020 2 repositories listedHere, we propose a new video instance segmentation framework built upon Transformers, termed VisTR, which views the VIS task as a direct end-to-end parallel sequence decoding/prediction problem.
-
24 May 2025 1 repository listedReasoning Video Object Segmentation is a challenging task, which generates a mask sequence from an input video and an implicit, complex text query.
-
1 Dec 2024 1 repository listedRecent DETR-based methods have advanced the development of Video Instance Segmentation (VIS) through transformers' efficiency and capability in modeling spatial and temporal information.
-
21 Sep 2024 1 repository listedThis way, while basing our method on an amodal instance segmentation, we nevertheless obtain video-level amodal instance segmentation results.
-
5 Sep 2024 1 repository listedTo this end, we introduce: (i) a new task termed \emph{space-time instance segmentation}, similar to video instance segmentation, whose goal is to segment instances throughout the entire duration of the sensor input…
-
29 Aug 2024 1 repository listedThis method is based on two key innovations: a Temporal Eigenvalue Loss (TEL) and a clip-level Quality Cluster Coefficient (QCC).
-
10 Jul 2024 1 repository listedWe discover that the domain gap between the VLM features (e.
-
3 Jul 2024 1 repository listedIn this paper, we introduce the Context-Aware Video Instance Segmentation (CAVIS), a novel framework designed to enhance instance association by integrating contextual information adjacent to each object.
-
28 Jun 2024 1 repository listedThis method achieves high video instance segmentation performance without manual video annotations, offering a cost-effective solution and new perspectives for video instance segmentation applications.
-
19 Mar 2024 1 repository listedThe experiments are performed on various video instance segmentation datasets, which demonstrate the effectiveness of our proposed method, especially for novel categories.
-
28 Feb 2024 1 repository listed Syntology ran 12 of 14 samples · 2 unverified · 14 pointer-only (licence)Despite the recent advances in unified image segmentation (IS), developing a unified video segmentation (VS) model remains a challenge.
-
14 Feb 2024 1 repository listedDeep video models, for example, 3D CNNs or video transformers, have achieved promising performance on sparse video tasks, i.
-
18 Jan 2024 1 repository listedTo mold instance queries to follow Brownian bridge and accomplish alignment with class texts, we design Bridge-Text Alignment (BTA) to learn discriminative bridge-level representations of instances via contrastive…
-
20 Dec 2023 1 repository listedWe present the \textbf{D}ecoupled \textbf{VI}deo \textbf{S}egmentation (DVIS) framework, a novel approach for the challenging task of universal video segmentation, including video instance segmentation (VIS), video…
Syntology lines on 9 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections