Papers › Tracking by Natural Language Specification

Tracking by Natural Language Specification

1 Jul 2017CVPR 2017 7archive 2025-07-28

Zhenyang Li, Ran Tao, Efstratios Gavves, Cees G. M. Snoek, Arnold W. M. Smeulders

This paper strives to track a target object in a video. Rather than specifying the target in the first frame of a video by a bounding box, we propose to track the object based on a natural language specification of the target, which provides a more natural human-machine interaction as well as a means to improve tracking results. We define three variants of tracking by language specification: one relying on lingual target specification only, one relying on visual target specification based on language, and one leveraging their joint capacity. To show the potential of tracking by natural language specification we extend two popular tracking datasets with lingual descriptions and report experiments. Finally, we also sketch new tracking scenarios in surveillance and other live video streams that become feasible with a lingual specification of the target.

PaperPDF

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Referring Expression Segmentation

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Referring Expression Segmentation A2D Sentences Li et al. AP 0.163 #21 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences Li et al. IoU mean 0.354 #21 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences Li et al. IoU overall 0.515 #21 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences Li et al. Precision@0.5 0.387 #21 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences Li et al. Precision@0.6 0.290 #21 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences Li et al. Precision@0.7 0.175 #21 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences Li et al. Precision@0.8 0.066 #21 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences Li et al. Precision@0.9 0.001 #21 of 27 Archive leaderboard report
Referring Expression Segmentation J-HMDB Li et al. AP 0.173 #17 of 21 Archive leaderboard report
Referring Expression Segmentation J-HMDB Li et al. IoU mean 0.491 #17 of 21 Archive leaderboard report
Referring Expression Segmentation J-HMDB Li et al. IoU overall 0.529 #17 of 21 Archive leaderboard report
Referring Expression Segmentation J-HMDB Li et al. Precision@0.5 0.578 #17 of 21 Archive leaderboard report
Referring Expression Segmentation J-HMDB Li et al. Precision@0.6 0.335 #17 of 21 Archive leaderboard report
Referring Expression Segmentation J-HMDB Li et al. Precision@0.7 0.103 #17 of 21 Archive leaderboard report
Referring Expression Segmentation J-HMDB Li et al. Precision@0.8 0.060 #17 of 21 Archive leaderboard report
Referring Expression Segmentation J-HMDB Li et al. Precision@0.9 0.000 #17 of 21 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections