Papers › Robust Online Video Instance Segmentation with Track Queries

Robust Online Video Instance Segmentation with Track Queries

16 Nov 2022arXiv:2211.09108archive 2025-07-28

Zitong Zhan, Daniel McKee, Svetlana Lazebnik

Recently, transformer-based methods have achieved impressive results on Video Instance Segmentation (VIS). However, most of these top-performing methods run in an offline manner by processing the entire video clip at once to predict instance mask volumes. This makes them incapable of handling the long videos that appear in challenging new video instance segmentation datasets like UVO and OVIS. We propose a fully online transformer-based video instance segmentation model that performs comparably to top offline methods on the YouTube-VIS 2019 benchmark and considerably outperforms them on UVO and OVIS. This method, called Robust Online Video Segmentation (ROVIS), augments the Mask2Former image instance segmentation model with track queries, a lightweight mechanism for carrying track information from frame to frame, originally introduced by the TrackFormer method for multi-object tracking. We show that, when combined with a strong enough image segmentation architecture, track queries can exhibit impressive accuracy while not being constrained to short videos.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

zitongzhan/mmtracking officialpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Image SegmentationInstance SegmentationMulti-Object TrackingObject TrackingSegmentationSemantic SegmentationVideo Instance SegmentationVideo SegmentationVideo Semantic Segmentation

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Video Instance Segmentation OVIS validation ROVIS (Swin-L) AP50 64.7 #17 of 44 Archive leaderboard report
Video Instance Segmentation OVIS validation ROVIS (Swin-L) AP75 42.6 #17 of 44 Archive leaderboard report
Video Instance Segmentation OVIS validation ROVIS (Swin-L) AR1 18.4 #17 of 44 Archive leaderboard report
Video Instance Segmentation OVIS validation ROVIS (Swin-L) AR10 49.1 #17 of 44 Archive leaderboard report
Video Instance Segmentation OVIS validation ROVIS (Swin-L) mask AP 42.6 #17 of 44 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

CLIP

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections