Papers › STC: Spatio-Temporal Contrastive Learning for Video Instance Segmentation

STC: Spatio-Temporal Contrastive Learning for Video Instance Segmentation

8 Feb 2022arXiv:2202.03747archive 2025-07-28

Zhengkai Jiang, Zhangxuan Gu, Jinlong Peng, Hang Zhou, Liang Liu, Yabiao Wang, Ying Tai, Chengjie Wang, Liqing Zhang

Video Instance Segmentation (VIS) is a task that simultaneously requires classification, segmentation, and instance association in a video. Recent VIS approaches rely on sophisticated pipelines to achieve this goal, including RoI-related operations or 3D convolutions. In contrast, we present a simple and efficient single-stage VIS framework based on the instance segmentation method CondInst by adding an extra tracking head. To improve instance association accuracy, a novel bi-directional spatio-temporal contrastive learning strategy for tracking embedding across frames is proposed. Moreover, an instance-wise temporal consistency scheme is utilized to produce temporally coherent results. Experiments conducted on the YouTube-VIS-2019, YouTube-VIS-2021, and OVIS-2021 datasets validate the effectiveness and efficiency of the proposed method. We hope the proposed framework can serve as a simple and strong alternative for many other instance-level video association tasks.

PaperPDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Contrastive LearningInstance SegmentationSegmentationSemantic SegmentationVideo Instance Segmentation

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Video Instance Segmentation OVIS validation STC (ResNet-50) AP50 33.5 #40 of 44 Archive leaderboard report
Video Instance Segmentation OVIS validation STC (ResNet-50) AP75 13.4 #40 of 44 Archive leaderboard report
Video Instance Segmentation OVIS validation STC (ResNet-50) mask AP 15.5 #40 of 44 Archive leaderboard report
Video Instance Segmentation YouTube-VIS validation STC (ResNet-50) AP50 57.2 #28 of 44 Archive leaderboard report
Video Instance Segmentation YouTube-VIS validation STC (ResNet-50) AP75 38.6 #28 of 44 Archive leaderboard report
Video Instance Segmentation YouTube-VIS validation STC (ResNet-50) AR1 36.9 #28 of 44 Archive leaderboard report
Video Instance Segmentation YouTube-VIS validation STC (ResNet-50) AR10 44.5 #28 of 44 Archive leaderboard report
Video Instance Segmentation YouTube-VIS validation STC (ResNet-50) mask AP 36.7 #28 of 44 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

CondInstContrastive Learning

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections