Papers › Actor and Action Modular Network for Text-based Video Segmentation

Actor and Action Modular Network for Text-based Video Segmentation

2 Nov 2020arXiv:2011.00786archive 2025-07-28

Jianhua Yang, Yan Huang, Kai Niu, Linjiang Huang, Zhanyu Ma, Liang Wang

Text-based video segmentation aims to segment an actor in video sequences by specifying the actor and its performing action with a textual query. Previous methods fail to explicitly align the video content with the textual query in a fine-grained manner according to the actor and its action, due to the problem of \emph{semantic asymmetry}. The \emph{semantic asymmetry} implies that two modalities contain different amounts of semantic information during the multi-modal fusion process. To alleviate this problem, we propose a novel actor and action modular network that individually localizes the actor and its action in two separate modules. Specifically, we first learn the actor-/action-related content from the video and textual query, and then match them in a symmetrical manner to localize the target tube. The target tube contains the desired actor and action which is then fed into a fully convolutional network to predict segmentation masks of the actor. Our method also establishes the association of objects cross multiple frames with the proposed temporal proposal aggregation mechanism. This enables our method to segment the video effectively and keep the temporal consistency of predictions. The whole model is allowed for joint learning of the actor-action matching and segmentation, as well as achieves the state-of-the-art performance for both single-frame segmentation and full video segmentation on A2D Sentences and J-HMDB Sentences datasets.

PaperPDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Action SegmentationAction UnderstandingReferring Expression SegmentationSegmentationSemantic SegmentationVideo SegmentationVideo Semantic Segmentation

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Referring Expression Segmentation A2D Sentences AAMN AP 0.396 #13 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences AAMN IoU mean 0.552 #13 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences AAMN IoU overall 0.617 #13 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences AAMN Precision@0.5 0.681 #13 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences AAMN Precision@0.6 0.629 #13 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences AAMN Precision@0.7 0.523 #13 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences AAMN Precision@0.8 0.296 #13 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences AAMN Precision@0.9 0.029 #13 of 27 Archive leaderboard report
Referring Expression Segmentation J-HMDB AAMN AP 0.321 #9 of 21 Archive leaderboard report
Referring Expression Segmentation J-HMDB AAMN IoU mean 0.576 #9 of 21 Archive leaderboard report
Referring Expression Segmentation J-HMDB AAMN IoU overall 0.583 #9 of 21 Archive leaderboard report
Referring Expression Segmentation J-HMDB AAMN Precision@0.5 0.773 #9 of 21 Archive leaderboard report
Referring Expression Segmentation J-HMDB AAMN Precision@0.6 0.627 #9 of 21 Archive leaderboard report
Referring Expression Segmentation J-HMDB AAMN Precision@0.7 0.360 #9 of 21 Archive leaderboard report
Referring Expression Segmentation J-HMDB AAMN Precision@0.8 0.044 #9 of 21 Archive leaderboard report
Referring Expression Segmentation J-HMDB AAMN Precision@0.9 0.000 #9 of 21 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections