Papers › Learning Motion-Appearance Co-Attention for Zero-Shot Video Object Segmentation
Learning Motion-Appearance Co-Attention for Zero-Shot Video Object Segmentation
Shu Yang, Lu Zhang, Jinqing Qi, Huchuan Lu, Shuo Wang, Xiaoxing Zhang
How to make the appearance and motion information interact effectively to accommodate complex scenarios is a fundamental issue in flow-based zero-shot video object segmentation. In this paper, we propose an Attentive Multi-Modality Collaboration Network (AMC-Net) to utilize appearance and motion information uniformly. Specifically, AMC-Net fuses robust information from multi-modality features and promotes their collaboration in two stages. First, we propose a Multi-Modality Co-Attention Gate (MCG) on the bilateral encoder branches, in which a gate function is used to formulate co-attention scores for balancing the contributions of multi-modality features and suppressing the redundant and misleading information. Then, we propose a Motion Correction Module (MCM) with a visual-motion attention mechanism, which is constructed to emphasize the features of foreground objects by incorporating the spatio-temporal correspondence between appearance and motion cues. Extensive experiments on three public challenging benchmark datasets verify that our proposed network performs favorably against existing state-of-the-art methods via training with fewer data.
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
1 archive task tag without a task page not shown.
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Unsupervised Video Object Segmentation | DAVIS 2016 val | AMC-Net | F | 84.6 | #14 of 25 | Archive leaderboard | report |
| Unsupervised Video Object Segmentation | DAVIS 2016 val | AMC-Net | G | 84.6 | #14 of 25 | Archive leaderboard | report |
| Unsupervised Video Object Segmentation | DAVIS 2016 val | AMC-Net | J | 84.5 | #14 of 25 | Archive leaderboard | report |
| Unsupervised Video Object Segmentation | FBMS test | AMC-Net | J | 76.5 | #12 of 15 | Archive leaderboard | report |
| Unsupervised Video Object Segmentation | YouTube-Objects | AMC-Net | J | 71.1 | #8 of 16 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections