Papers › Asymmetric Cross-Guided Attention Network for Actor and Action Video Segmentation From...

Asymmetric Cross-Guided Attention Network for Actor and Action Video Segmentation From Natural Language Query

1 Oct 2019ICCV 2019 10archive 2025-07-28

Hao Wang, Cheng Deng, Junchi Yan, Dacheng Tao

Actor and action video segmentation from natural language query aims to selectively segment the actor and its action in a video based on an input textual description. Previous works mostly focus on learning simple correlation between two heterogeneous features of vision and language via dynamic convolution or fully convolutional classification. However, they ignore the linguistic variation of natural language query and have difficulty in modeling global visual context, which leads to unsatisfactory segmentation performance. To address these issues, we propose an asymmetric cross-guided attention network for actor and action video segmentation from natural language query. Specifically, we frame an asymmetric cross-guided attention network, which consists of vision guided language attention to reduce the linguistic variation of input query and language guided vision attention to incorporate query-focused global visual context simultaneously. Moreover, we adopt multi-resolution fusion scheme and weighted loss for foreground and background pixels to obtain further performance improvement. Extensive experiments on Actor-Action Dataset Sentences and J-HMDB Sentences show that our proposed approach notably outperforms state-of-the-art methods.

PaperPDFCode

Code

haowang1992/ACGA officialpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Referring Expression SegmentationSegmentationVideo SegmentationVideo Semantic Segmentation

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Referring Expression Segmentation A2D Sentences ACGA AP 0.274 #18 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences ACGA IoU mean 0.490 #18 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences ACGA IoU overall 0.601 #18 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences ACGA Precision@0.5 0.557 #18 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences ACGA Precision@0.6 0.459 #18 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences ACGA Precision@0.7 0.319 #18 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences ACGA Precision@0.8 0.16 #18 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences ACGA Precision@0.9 0.02 #18 of 27 Archive leaderboard report
Referring Expression Segmentation J-HMDB ACGA AP 0.289 #12 of 21 Archive leaderboard report
Referring Expression Segmentation J-HMDB ACGA IoU mean 0.584 #12 of 21 Archive leaderboard report
Referring Expression Segmentation J-HMDB ACGA IoU overall 0.576 #12 of 21 Archive leaderboard report
Referring Expression Segmentation J-HMDB ACGA Precision@0.5 0.756 #12 of 21 Archive leaderboard report
Referring Expression Segmentation J-HMDB ACGA Precision@0.6 0.564 #12 of 21 Archive leaderboard report
Referring Expression Segmentation J-HMDB ACGA Precision@0.7 0.287 #12 of 21 Archive leaderboard report
Referring Expression Segmentation J-HMDB ACGA Precision@0.8 0.034 #12 of 21 Archive leaderboard report
Referring Expression Segmentation J-HMDB ACGA Precision@0.9 0.000 #12 of 21 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Convolution

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections