Papers › STC: Spatio-Temporal Contrastive Learning for Video Instance Segmentation
STC: Spatio-Temporal Contrastive Learning for Video Instance Segmentation
Zhengkai Jiang, Zhangxuan Gu, Jinlong Peng, Hang Zhou, Liang Liu, Yabiao Wang, Ying Tai, Chengjie Wang, Liqing Zhang
Video Instance Segmentation (VIS) is a task that simultaneously requires classification, segmentation, and instance association in a video. Recent VIS approaches rely on sophisticated pipelines to achieve this goal, including RoI-related operations or 3D convolutions. In contrast, we present a simple and efficient single-stage VIS framework based on the instance segmentation method CondInst by adding an extra tracking head. To improve instance association accuracy, a novel bi-directional spatio-temporal contrastive learning strategy for tracking embedding across frames is proposed. Moreover, an instance-wise temporal consistency scheme is utilized to produce temporally coherent results. Experiments conducted on the YouTube-VIS-2019, YouTube-VIS-2021, and OVIS-2021 datasets validate the effectiveness and efficiency of the proposed method. We hope the proposed framework can serve as a simple and strong alternative for many other instance-level video association tasks.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
No code repository is listed for this paper in the archive or in Syntology's graph.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Video Instance Segmentation | OVIS validation | STC (ResNet-50) | AP50 | 33.5 | #40 of 44 | Archive leaderboard | report |
| Video Instance Segmentation | OVIS validation | STC (ResNet-50) | AP75 | 13.4 | #40 of 44 | Archive leaderboard | report |
| Video Instance Segmentation | OVIS validation | STC (ResNet-50) | mask AP | 15.5 | #40 of 44 | Archive leaderboard | report |
| Video Instance Segmentation | YouTube-VIS validation | STC (ResNet-50) | AP50 | 57.2 | #28 of 44 | Archive leaderboard | report |
| Video Instance Segmentation | YouTube-VIS validation | STC (ResNet-50) | AP75 | 38.6 | #28 of 44 | Archive leaderboard | report |
| Video Instance Segmentation | YouTube-VIS validation | STC (ResNet-50) | AR1 | 36.9 | #28 of 44 | Archive leaderboard | report |
| Video Instance Segmentation | YouTube-VIS validation | STC (ResNet-50) | AR10 | 44.5 | #28 of 44 | Archive leaderboard | report |
| Video Instance Segmentation | YouTube-VIS validation | STC (ResNet-50) | mask AP | 36.7 | #28 of 44 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Methods
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections