Papers › Context Modulated Dynamic Networks for Actor and Action Video Segmentation with...

Context Modulated Dynamic Networks for Actor and Action Video Segmentation with Language Queries

3 Apr 2020archive 2025-07-28

Hao Wang, Cheng Deng, Fan Ma, Yi Yang

Actor and action video segmentation with language queries aims to segment out the expression referred objects in the video. This process requires comprehensive language reasoning and fine-grained video understanding. Previous methods mainly leverage dynamic convolutional networks to match visual and semantic representations. However, the dynamic convolution neglects spatial context when processing each region in the frame and is thus challenging to segment similar objects in the complex scenarios. To address such limitation, we construct a context modulated dynamic convolutional network. Specifically, we propose a context modulated dynamic convolutional operation in the proposed framework. The kernels for the specific region are generated from both language sentences and surrounding context features. Moreover, we devise a temporal encoder to incorporate motions into the visual features to further match the query descriptions. Extensive experiments on two benchmark datasets, Actor-Action Dataset Sentences (A2D Sentences) and J-HMDB Sentences, demonstrate that our proposed approach notably outperforms state-of-the-art methods.

PaperPDF

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Referring Expression SegmentationVideo SegmentationVideo Semantic SegmentationVideo Understanding

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Referring Expression Segmentation A2D Sentences CMDy AP 0.333 #16 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences CMDy IoU mean 0.531 #16 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences CMDy IoU overall 0.623 #16 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences CMDy Precision@0.5 0.607 #16 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences CMDy Precision@0.6 0.525 #16 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences CMDy Precision@0.7 0.405 #16 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences CMDy Precision@0.8 0.235 #16 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences CMDy Precision@0.9 0.045 #16 of 27 Archive leaderboard report
Referring Expression Segmentation J-HMDB CMDy AP 0.301 #10 of 21 Archive leaderboard report
Referring Expression Segmentation J-HMDB CMDy IoU mean 0.576 #10 of 21 Archive leaderboard report
Referring Expression Segmentation J-HMDB CMDy IoU overall 0.554 #10 of 21 Archive leaderboard report
Referring Expression Segmentation J-HMDB CMDy Precision@0.5 0.742 #10 of 21 Archive leaderboard report
Referring Expression Segmentation J-HMDB CMDy Precision@0.6 0.587 #10 of 21 Archive leaderboard report
Referring Expression Segmentation J-HMDB CMDy Precision@0.7 0.316 #10 of 21 Archive leaderboard report
Referring Expression Segmentation J-HMDB CMDy Precision@0.8 0.047 #10 of 21 Archive leaderboard report
Referring Expression Segmentation J-HMDB CMDy Precision@0.9 0.000 #10 of 21 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections