Papers › Occluded Video Instance Segmentation: A Benchmark

Occluded Video Instance Segmentation: A Benchmark

2 Feb 2021arXiv:2102.01558archive 2025-07-28

Jiyang Qi, Yan Gao, Yao Hu, Xinggang Wang, Xiaoyu Liu, Xiang Bai, Serge Belongie, Alan Yuille, Philip H. S. Torr, Song Bai

Can our video understanding systems perceive objects when a heavy occlusion exists in a scene? To answer this question, we collect a large-scale dataset called OVIS for occluded video instance segmentation, that is, to simultaneously detect, segment, and track instances in occluded scenes. OVIS consists of 296k high-quality instance masks from 25 semantic categories, where object occlusions usually occur. While our human vision systems can understand those occluded instances by contextual reasoning and association, our experiments suggest that current video understanding systems cannot. On the OVIS dataset, the highest AP achieved by state-of-the-art algorithms is only 16.3, which reveals that we are still at a nascent stage for understanding objects, instances, and videos in a real-world scenario. We also present a simple plug-and-play module that performs temporal feature calibration to complement missing object cues caused by occlusion. Built upon MaskTrack R-CNN and SipMask, we obtain a remarkable AP improvement on the OVIS dataset. The OVIS dataset and project code are available at http://songbai.site/ovis .

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

haochenheheda/lvvis mentioned on GitHubpytorchGPL-3.0 report
qjy981010/CMaskTrack-RCNN mentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Instance SegmentationSegmentationSemantic SegmentationVideo Instance SegmentationVideo Understanding

Datasets

Introduced by this paper, per the archive.

OVIS

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Video Instance Segmentation OVIS validation CMaskTrack R-CNN (ResNet-50) AP50 33.9 #41 of 44 Archive leaderboard report
Video Instance Segmentation OVIS validation CMaskTrack R-CNN (ResNet-50) AP75 13.1 #41 of 44 Archive leaderboard report
Video Instance Segmentation OVIS validation CMaskTrack R-CNN (ResNet-50) APho 4.1 #41 of 44 Archive leaderboard report
Video Instance Segmentation OVIS validation CMaskTrack R-CNN (ResNet-50) APmo 18.7 #41 of 44 Archive leaderboard report
Video Instance Segmentation OVIS validation CMaskTrack R-CNN (ResNet-50) APso 28.6 #41 of 44 Archive leaderboard report
Video Instance Segmentation OVIS validation CMaskTrack R-CNN (ResNet-50) mask AP 15.4 #41 of 44 Archive leaderboard report
Video Instance Segmentation OVIS validation CSipMask (ResNet-50) AP50 29.9 #44 of 44 Archive leaderboard report
Video Instance Segmentation OVIS validation CSipMask (ResNet-50) AP75 12.5 #44 of 44 Archive leaderboard report
Video Instance Segmentation OVIS validation CSipMask (ResNet-50) APho 2.7 #44 of 44 Archive leaderboard report
Video Instance Segmentation OVIS validation CSipMask (ResNet-50) APmo 12.8 #44 of 44 Archive leaderboard report
Video Instance Segmentation OVIS validation CSipMask (ResNet-50) APso 23 #44 of 44 Archive leaderboard report
Video Instance Segmentation OVIS validation CSipMask (ResNet-50) mask AP 14.3 #44 of 44 Archive leaderboard report
Video Instance Segmentation YouTube-VIS validation CSipMask AP50 55.6 #34 of 44 Archive leaderboard report
Video Instance Segmentation YouTube-VIS validation CSipMask AP75 38.1 #34 of 44 Archive leaderboard report
Video Instance Segmentation YouTube-VIS validation CSipMask mask AP 35.1 #34 of 44 Archive leaderboard report
Video Instance Segmentation YouTube-VIS validation CMaskTrack R-CNN AP50 52.8 #39 of 44 Archive leaderboard report
Video Instance Segmentation YouTube-VIS validation CMaskTrack R-CNN AP75 34.9 #39 of 44 Archive leaderboard report
Video Instance Segmentation YouTube-VIS validation CMaskTrack R-CNN mask AP 32.1 #39 of 44 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections