Papers › DVIS: Decoupled Video Instance Segmentation Framework

DVIS: Decoupled Video Instance Segmentation Framework

6 Jun 2023ICCV 2023 1arXiv:2306.03413archive 2025-07-28

Tao Zhang, Xingye Tian, Yu Wu, Shunping Ji, Xuebo Wang, Yuan Zhang, Pengfei Wan

Video instance segmentation (VIS) is a critical task with diverse applications, including autonomous driving and video editing. Existing methods often underperform on complex and long videos in real world, primarily due to two factors. Firstly, offline methods are limited by the tightly-coupled modeling paradigm, which treats all frames equally and disregards the interdependencies between adjacent frames. Consequently, this leads to the introduction of excessive noise during long-term temporal alignment. Secondly, online methods suffer from inadequate utilization of temporal information. To tackle these challenges, we propose a decoupling strategy for VIS by dividing it into three independent sub-tasks: segmentation, tracking, and refinement. The efficacy of the decoupling strategy relies on two crucial elements: 1) attaining precise long-term alignment outcomes via frame-by-frame association during tracking, and 2) the effective utilization of temporal information predicated on the aforementioned accurate alignment outcomes during refinement. We introduce a novel referring tracker and temporal refiner to construct the \textbf{D}ecoupled \textbf{VIS} framework (\textbf{DVIS}). DVIS achieves new SOTA performance in both VIS and VPS, surpassing the current SOTA methods by 7.3 AP and 9.6 VPQ on the OVIS and VIPSeg datasets, which are the most challenging and realistic benchmarks. Moreover, thanks to the decoupling strategy, the referring tracker and temporal refiner are super light-weight (only 1.69\% of the segmenter FLOPs), allowing for efficient training and inference on a single GPU with 11G memory. The code is available at \href{https://github.com/zhang-tao-whu/DVIS}{https://github.com/zhang-tao-whu/DVIS}.

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

zhang-tao-whu/DVIS officialmentioned in papermentioned on GitHubpytorchMIT report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Autonomous DrivingInstance SegmentationSegmentationSemantic SegmentationVideo EditingVideo Instance SegmentationVideo Panoptic Segmentation

1 archive task tag without a task page not shown.

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Video Instance Segmentation OVIS validation DVIS(Swin-L, Offline) AP50 75.9 #5 of 44 Archive leaderboard report
Video Instance Segmentation OVIS validation DVIS(Swin-L, Offline) AP75 53.0 #5 of 44 Archive leaderboard report
Video Instance Segmentation OVIS validation DVIS(Swin-L, Offline) AR1 19.4 #5 of 44 Archive leaderboard report
Video Instance Segmentation OVIS validation DVIS(Swin-L, Offline) AR10 55.3 #5 of 44 Archive leaderboard report
Video Instance Segmentation OVIS validation DVIS(Swin-L, Offline) mask AP 49.9 #5 of 44 Archive leaderboard report
Video Instance Segmentation OVIS validation DVIS(Swin-L, Online) AP50 71.9 #8 of 44 Archive leaderboard report
Video Instance Segmentation OVIS validation DVIS(Swin-L, Online) AP75 49.2 #8 of 44 Archive leaderboard report
Video Instance Segmentation OVIS validation DVIS(Swin-L, Online) AR1 19.4 #8 of 44 Archive leaderboard report
Video Instance Segmentation OVIS validation DVIS(Swin-L, Online) AR10 52.5 #8 of 44 Archive leaderboard report
Video Instance Segmentation OVIS validation DVIS(Swin-L, Online) mask AP 47.1 #8 of 44 Archive leaderboard report
Video Instance Segmentation YouTube-VIS 2021 DVIS(Swin-L) AP50 83.0 #8 of 26 Archive leaderboard report
Video Instance Segmentation YouTube-VIS 2021 DVIS(Swin-L) AP75 68.4 #8 of 26 Archive leaderboard report
Video Instance Segmentation YouTube-VIS 2021 DVIS(Swin-L) AR1 47.7 #8 of 26 Archive leaderboard report
Video Instance Segmentation YouTube-VIS 2021 DVIS(Swin-L) AR10 65.7 #8 of 26 Archive leaderboard report
Video Instance Segmentation YouTube-VIS 2021 DVIS(Swin-L) mask AP 60.1 #8 of 26 Archive leaderboard report
Video Instance Segmentation YouTube-VIS validation DVIS AP50 88.0 #3 of 44 Archive leaderboard report
Video Instance Segmentation YouTube-VIS validation DVIS AP75 72.7 #3 of 44 Archive leaderboard report
Video Instance Segmentation YouTube-VIS validation DVIS AR1 56.5 #3 of 44 Archive leaderboard report
Video Instance Segmentation YouTube-VIS validation DVIS AR10 70.3 #3 of 44 Archive leaderboard report
Video Instance Segmentation YouTube-VIS validation DVIS mask AP 64.9 #3 of 44 Archive leaderboard report
Video Instance Segmentation Youtube-VIS 2022 Validation DVIS(Swin-L) AP50_L 69.0 #4 of 7 Archive leaderboard report
Video Instance Segmentation Youtube-VIS 2022 Validation DVIS(Swin-L) AP75_L 48.8 #4 of 7 Archive leaderboard report
Video Instance Segmentation Youtube-VIS 2022 Validation DVIS(Swin-L) AR10_L 51.8 #4 of 7 Archive leaderboard report
Video Instance Segmentation Youtube-VIS 2022 Validation DVIS(Swin-L) AR1_L 37.2 #4 of 7 Archive leaderboard report
Video Instance Segmentation Youtube-VIS 2022 Validation DVIS(Swin-L) mAP_L 45.9 #4 of 7 Archive leaderboard report
Video Panoptic Segmentation VIPSeg DVIS(Swin-L) STQ 55.3 #4 of 12 Archive leaderboard report
Video Panoptic Segmentation VIPSeg DVIS(Swin-L) VPQ 57.6 #4 of 12 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections