Papers › Polar Relative Positional Encoding for Video-Language Segmentation

Polar Relative Positional Encoding for Video-Language Segmentation

20 Jul 2020archive 2025-07-28

Ke Ning, Lingxi Xie, Fei Wu, Qi Tian

In this paper, we tackle a challenging task named video-language segmentation. Given a video and a sentence in natural language, the goal is to segment the object or actor described by the sentence in video frames. To accurately denote a target object, the given sentence usually refers to multiple attributes, such as nearby objects with spatial relations, etc. In this paper, we propose a novel Polar Relative Positional Encoding (PRPE) mechanism that represents spatial relations in a ``linguistic'' way, i.e., in terms of direction and range. Sentence feature can interact with positional embeddings in a more direct way to extract the implied relative positional relations. We also propose parameterized functions for these positional embeddings to adapt real-value directions and ranges. With PRPE, we design a Polar Attention Module (PAM) as the basic module for vision-language fusion. Our method outperforms previous best method by a large margin of 11.4% absolute improvement in terms of mAP on the challenging A2D Sentences dataset. Our method also achieves competitive performances on the J-HMDB Sentences dataset.

PaperPDF

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Referring Expression SegmentationSentence

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Referring Expression Segmentation A2D Sentences PRPE AP 0.388 #14 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences PRPE IoU mean 0.529 #14 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences PRPE IoU overall 0.661 #14 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences PRPE Precision@0.5 0.634 #14 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences PRPE Precision@0.6 0.579 #14 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences PRPE Precision@0.7 0.483 #14 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences PRPE Precision@0.8 0.322 #14 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences PRPE Precision@0.9 0.083 #14 of 27 Archive leaderboard report
Referring Expression Segmentation J-HMDB PRPE AP 0.294 #11 of 21 Archive leaderboard report
Referring Expression Segmentation J-HMDB PRPE Precision@0.5 0.572 #11 of 21 Archive leaderboard report
Referring Expression Segmentation J-HMDB PRPE Precision@0.6 0.690 #11 of 21 Archive leaderboard report
Referring Expression Segmentation J-HMDB PRPE Precision@0.7 0.319 #11 of 21 Archive leaderboard report
Referring Expression Segmentation J-HMDB PRPE Precision@0.8 0.06 #11 of 21 Archive leaderboard report
Referring Expression Segmentation J-HMDB PRPE Precision@0.9 0.001 #11 of 21 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections