Papers › 1st Place Solution for YouTubeVOS Challenge 2021:Video Instance Segmentation
1st Place Solution for YouTubeVOS Challenge 2021:Video Instance Segmentation
Thuy C. Nguyen, Tuan N. Tang, Nam LH. Phan, Chuong H. Nguyen, Masayuki Yamazaki, Masao Yamanaka
Video Instance Segmentation (VIS) is a multi-task problem performing detection, segmentation, and tracking simultaneously. Extended from image set applications, video data additionally induces the temporal information, which, if handled appropriately, is very useful to identify and predict object motions. In this work, we design a unified model to mutually learn these tasks. Specifically, we propose two modules, named Temporally Correlated Instance Segmentation (TCIS) and Bidirectional Tracking (BiTrack), to take the benefit of the temporal correlation between the object's instance masks across adjacent frames. On the other hand, video data is often redundant due to the frame's overlap. Our analysis shows that this problem is particularly severe for the YoutubeVOS-VIS2021 data. Therefore, we propose a Multi-Source Data (MSD) training mechanism to compensate for the data deficiency. By combining these techniques with a bag of tricks, the network performance is significantly boosted compared to the baseline, and outperforms other methods by a considerable margin on the YoutubeVOS-VIS 2019 and 2021 datasets.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
No code repository is listed for this paper in the archive or in Syntology's graph.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Video Instance Segmentation | YouTube-VIS validation | TCIS (Swin-S) | AP50 | 76.6 | #12 of 44 | Archive leaderboard | report |
| Video Instance Segmentation | YouTube-VIS validation | TCIS (Swin-S) | AP75 | 65.6 | #12 of 44 | Archive leaderboard | report |
| Video Instance Segmentation | YouTube-VIS validation | TCIS (Swin-S) | AR1 | 47 | #12 of 44 | Archive leaderboard | report |
| Video Instance Segmentation | YouTube-VIS validation | TCIS (Swin-S) | AR10 | 57.9 | #12 of 44 | Archive leaderboard | report |
| Video Instance Segmentation | YouTube-VIS validation | TCIS (Swin-S) | mask AP | 54.3 | #12 of 44 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections