Papers › InstanceFormer: An Online Video Instance Segmentation Framework

InstanceFormer: An Online Video Instance Segmentation Framework

22 Aug 2022arXiv:2208.10547archive 2025-07-28

Rajat Koner, Tanveer Hannan, Suprosanna Shit, Sahand Sharifzadeh, Matthias Schubert, Thomas Seidl, Volker Tresp

Recent transformer-based offline video instance segmentation (VIS) approaches achieve encouraging results and significantly outperform online approaches. However, their reliance on the whole video and the immense computational complexity caused by full Spatio-temporal attention limit them in real-life applications such as processing lengthy videos. In this paper, we propose a single-stage transformer-based efficient online VIS framework named InstanceFormer, which is especially suitable for long and challenging videos. We propose three novel components to model short-term and long-term dependency and temporal coherence. First, we propagate the representation, location, and semantic information of prior instances to model short-term changes. Second, we propose a novel memory cross-attention in the decoder, which allows the network to look into earlier instances within a certain temporal window. Finally, we employ a temporal contrastive loss to impose coherence in the representation of an instance across all frames. Memory attention and temporal coherence are particularly beneficial to long-range dependency modeling, including challenging scenarios like occlusion. The proposed InstanceFormer outperforms previous online benchmark methods by a large margin across multiple datasets. Most importantly, InstanceFormer surpasses offline approaches for challenging and long datasets such as YouTube-VIS-2021 and OVIS. Code is available at https://github.com/rajatkoner08/InstanceFormer.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

rajatkoner08/instanceformer officialmentioned in papermentioned on GitHubpytorchApache-2.0 report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

DecoderInstance SegmentationSemantic SegmentationVideo Instance Segmentation

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Video Instance Segmentation OVIS validation InstanceFormer (Swin-L) AP50 42.5 #34 of 44 Archive leaderboard report
Video Instance Segmentation OVIS validation InstanceFormer (Swin-L) AP75 21.61 #34 of 44 Archive leaderboard report
Video Instance Segmentation OVIS validation InstanceFormer (Swin-L) AR1 12.9 #34 of 44 Archive leaderboard report
Video Instance Segmentation OVIS validation InstanceFormer (Swin-L) AR10 29.3 #34 of 44 Archive leaderboard report
Video Instance Segmentation OVIS validation InstanceFormer (Swin-L) mask AP 22.8 #34 of 44 Archive leaderboard report
Video Instance Segmentation OVIS validation InstanceFormer(ResNet-50) AP50 40.7 #35 of 44 Archive leaderboard report
Video Instance Segmentation OVIS validation InstanceFormer(ResNet-50) AP75 18.1 #35 of 44 Archive leaderboard report
Video Instance Segmentation OVIS validation InstanceFormer(ResNet-50) AR1 12 #35 of 44 Archive leaderboard report
Video Instance Segmentation OVIS validation InstanceFormer(ResNet-50) AR10 27.1 #35 of 44 Archive leaderboard report
Video Instance Segmentation OVIS validation InstanceFormer(ResNet-50) mask AP 20.0 #35 of 44 Archive leaderboard report
Video Instance Segmentation YouTube-VIS 2021 InstanceFormer (Swin-L) AP50 73.7 #19 of 26 Archive leaderboard report
Video Instance Segmentation YouTube-VIS 2021 InstanceFormer (Swin-L) AP75 56.9 #19 of 26 Archive leaderboard report
Video Instance Segmentation YouTube-VIS 2021 InstanceFormer (Swin-L) AR1 42.8 #19 of 26 Archive leaderboard report
Video Instance Segmentation YouTube-VIS 2021 InstanceFormer (Swin-L) AR10 56.0 #19 of 26 Archive leaderboard report
Video Instance Segmentation YouTube-VIS 2021 InstanceFormer (Swin-L) mask AP 51.0 #19 of 26 Archive leaderboard report
Video Instance Segmentation YouTube-VIS 2021 InstanceFormer (ResNet-50) AP50 62.4 #25 of 26 Archive leaderboard report
Video Instance Segmentation YouTube-VIS 2021 InstanceFormer (ResNet-50) AP75 43.7 #25 of 26 Archive leaderboard report
Video Instance Segmentation YouTube-VIS 2021 InstanceFormer (ResNet-50) AR1 36.1 #25 of 26 Archive leaderboard report
Video Instance Segmentation YouTube-VIS 2021 InstanceFormer (ResNet-50) AR10 48.1 #25 of 26 Archive leaderboard report
Video Instance Segmentation YouTube-VIS 2021 InstanceFormer (ResNet-50) mask AP 40.8 #25 of 26 Archive leaderboard report
Video Instance Segmentation YouTube-VIS validation InstanceFormer(Swin-L) AP50 78.0 #11 of 44 Archive leaderboard report
Video Instance Segmentation YouTube-VIS validation InstanceFormer(Swin-L) AP75 64.2 #11 of 44 Archive leaderboard report
Video Instance Segmentation YouTube-VIS validation InstanceFormer(Swin-L) AR1 50.9 #11 of 44 Archive leaderboard report
Video Instance Segmentation YouTube-VIS validation InstanceFormer(Swin-L) AR10 61.6 #11 of 44 Archive leaderboard report
Video Instance Segmentation YouTube-VIS validation InstanceFormer(Swin-L) mask AP 56.3 #11 of 44 Archive leaderboard report
Video Instance Segmentation YouTube-VIS validation InstanceFormer(ResNet-50) AP50 68.6 #21 of 44 Archive leaderboard report
Video Instance Segmentation YouTube-VIS validation InstanceFormer(ResNet-50) AP75 49.6 #21 of 44 Archive leaderboard report
Video Instance Segmentation YouTube-VIS validation InstanceFormer(ResNet-50) AR1 42.1 #21 of 44 Archive leaderboard report
Video Instance Segmentation YouTube-VIS validation InstanceFormer(ResNet-50) AR10 53.5 #21 of 44 Archive leaderboard report
Video Instance Segmentation YouTube-VIS validation InstanceFormer(ResNet-50) mask AP 45.6 #21 of 44 Archive leaderboard report
Video Instance Segmentation Youtube-VIS 2022 Validation InstanceFormer (Swin) AP50_L 44.6 #6 of 7 Archive leaderboard report
Video Instance Segmentation Youtube-VIS 2022 Validation InstanceFormer (Swin) AP75_L 27.3 #6 of 7 Archive leaderboard report
Video Instance Segmentation Youtube-VIS 2022 Validation InstanceFormer (Swin) AR10_L 29.2 #6 of 7 Archive leaderboard report
Video Instance Segmentation Youtube-VIS 2022 Validation InstanceFormer (Swin) AR1_L 25.0 #6 of 7 Archive leaderboard report
Video Instance Segmentation Youtube-VIS 2022 Validation InstanceFormer (Swin) mAP_L 26.3 #6 of 7 Archive leaderboard report
Video Instance Segmentation Youtube-VIS 2022 Validation InstanceFormer (Resnet-50) AP50_L 49.5 #7 of 7 Archive leaderboard report
Video Instance Segmentation Youtube-VIS 2022 Validation InstanceFormer (Resnet-50) AP75_L 26.7 #7 of 7 Archive leaderboard report
Video Instance Segmentation Youtube-VIS 2022 Validation InstanceFormer (Resnet-50) AR10_L 30.1 #7 of 7 Archive leaderboard report
Video Instance Segmentation Youtube-VIS 2022 Validation InstanceFormer (Resnet-50) AR1_L 23.9 #7 of 7 Archive leaderboard report
Video Instance Segmentation Youtube-VIS 2022 Validation InstanceFormer (Resnet-50) mAP_L 24.8 #7 of 7 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections