Papers › Fast Video Object Segmentation via Dynamic Targeting Network
Fast Video Object Segmentation via Dynamic Targeting Network
Lu Zhang, Zhe Lin, Jianming Zhang, Huchuan Lu, You He
We propose a new model for fast and accurate video object segmentation. It consists of two convolutional neural networks, a Dynamic Targeting Network (DTN) and a Mask Refinement Network (MRN). DTN locates the object by dynamically focusing on regions of interest surrounding the target object. The target region is predicted by DTN via two sub-streams, Box Propagation (BP) and Box Re-identification (BR). The BP stream is faster but less effective at objects with large deformation or occlusion. The BR stream performs better in difficult scenarios at a higher computation cost. We propose a Decision Module (DM) to adaptively determine which sub-stream to use for each frame. Finally, MRN is exploited to predict segmentation within the target region. Experimental results on two public datasets demonstrate that the proposed model significantly outperforms existing methods without online training in both accuracy and efficiency, and is comparable to online training-based methods in accuracy with an order of magnitude faster speed.
Code
No code repository is listed for this paper in the archive or in Syntology's graph.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Semi-Supervised Video Object Segmentation | DAVIS (no YouTube-VOS training) | DTN | D16 val (F) | 83.5 | #23 of 26 | Archive leaderboard | report |
| Semi-Supervised Video Object Segmentation | DAVIS (no YouTube-VOS training) | DTN | D16 val (G) | 83.6 | #23 of 26 | Archive leaderboard | report |
| Semi-Supervised Video Object Segmentation | DAVIS (no YouTube-VOS training) | DTN | D16 val (J) | 83.7 | #23 of 26 | Archive leaderboard | report |
| Semi-Supervised Video Object Segmentation | DAVIS (no YouTube-VOS training) | DTN | D17 val (F) | 70.6 | #23 of 26 | Archive leaderboard | report |
| Semi-Supervised Video Object Segmentation | DAVIS (no YouTube-VOS training) | DTN | D17 val (G) | 67.4 | #23 of 26 | Archive leaderboard | report |
| Semi-Supervised Video Object Segmentation | DAVIS (no YouTube-VOS training) | DTN | D17 val (J) | 64.2 | #23 of 26 | Archive leaderboard | report |
| Semi-Supervised Video Object Segmentation | DAVIS (no YouTube-VOS training) | DTN | FPS | 14.3 | #23 of 26 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections