Papers › Crossover Learning for Fast Online Video Instance Segmentation

Crossover Learning for Fast Online Video Instance Segmentation

13 Apr 2021ICCV 2021 10arXiv:2104.05970archive 2025-07-28

Shusheng Yang, Yuxin Fang, Xinggang Wang, Yu Li, Chen Fang, Ying Shan, Bin Feng, Wenyu Liu

Modeling temporal visual context across frames is critical for video instance segmentation (VIS) and other video understanding tasks. In this paper, we propose a fast online VIS model named CrossVIS. For temporal information modeling in VIS, we present a novel crossover learning scheme that uses the instance feature in the current frame to pixel-wisely localize the same instance in other frames. Different from previous schemes, crossover learning does not require any additional network parameters for feature enhancement. By integrating with the instance segmentation loss, crossover learning enables efficient cross-frame instance-to-pixel relation learning and brings cost-free improvement during inference. Besides, a global balanced instance embedding branch is proposed for more accurate and more stable online instance association. We conduct extensive experiments on three challenging VIS benchmarks, \ie, YouTube-VIS-2019, OVIS, and YouTube-VIS-2021 to evaluate our methods. To our knowledge, CrossVIS achieves state-of-the-art performance among all online VIS methods and shows a decent trade-off between latency and accuracy. Code will be available to facilitate future research.

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

hustvl/CrossVIS officialpytorchNOASSERTION report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Instance SegmentationSemantic SegmentationVideo Instance SegmentationVideo Understanding

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Video Instance Segmentation OVIS validation CrossVIS (ResNet-50, calibration) AP50 35.5 #36 of 44 Archive leaderboard report
Video Instance Segmentation OVIS validation CrossVIS (ResNet-50, calibration) AP75 16.9 #36 of 44 Archive leaderboard report
Video Instance Segmentation OVIS validation CrossVIS (ResNet-50, calibration) mask AP 18.1 #36 of 44 Archive leaderboard report
Video Instance Segmentation OVIS validation CrossVIS (ResNet-50) AP50 32.7 #43 of 44 Archive leaderboard report
Video Instance Segmentation OVIS validation CrossVIS (ResNet-50) AP75 12.1 #43 of 44 Archive leaderboard report
Video Instance Segmentation OVIS validation CrossVIS (ResNet-50) mask AP 14.9 #43 of 44 Archive leaderboard report
Video Instance Segmentation YouTube-VIS validation CrossVIS (ResNet-101) AP50 57.3 #29 of 44 Archive leaderboard report
Video Instance Segmentation YouTube-VIS validation CrossVIS (ResNet-101) AP75 39.7 #29 of 44 Archive leaderboard report
Video Instance Segmentation YouTube-VIS validation CrossVIS (ResNet-101) AR1 36 #29 of 44 Archive leaderboard report
Video Instance Segmentation YouTube-VIS validation CrossVIS (ResNet-101) AR10 42 #29 of 44 Archive leaderboard report
Video Instance Segmentation YouTube-VIS validation CrossVIS (ResNet-101) mask AP 36.6 #29 of 44 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections