Papers › Blazingly Fast Video Object Segmentation with Pixel-Wise Metric Learning

Blazingly Fast Video Object Segmentation with Pixel-Wise Metric Learning

9 Apr 2018CVPR 2018 6arXiv:1804.03131archive 2025-07-28

Yuhua Chen, Jordi Pont-Tuset, Alberto Montes, Luc van Gool

This paper tackles the problem of video object segmentation, given some user annotation which indicates the object of interest. The problem is formulated as pixel-wise retrieval in a learned embedding space: we embed pixels of the same object instance into the vicinity of each other, using a fully convolutional network trained by a modified triplet loss as the embedding model. Then the annotated pixels are set as reference and the rest of the pixels are classified using a nearest-neighbor approach. The proposed method supports different kinds of user input such as segmentation mask in the first frame (semi-supervised scenario), or a sparse set of clicked points (interactive scenario). In the semi-supervised scenario, we achieve results competitive with the state of the art but at a fraction of computation cost (275 milliseconds per frame). In the interactive scenario where the user is able to refine their input iteratively, the proposed method provides instant response to each input, and reaches comparable quality to competing methods with much less interaction.

PaperPDFConference PDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Metric LearningObjectRetrievalSegmentationSemantic SegmentationSemi-Supervised Video Object SegmentationVideo Object SegmentationVideo Semantic SegmentationVisual Object Tracking

1 archive task tag without a task page not shown.

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Semi-Supervised Video Object Segmentation DAVIS 2016 PML F-measure (Decay) 7.8 #66 of 78 Archive leaderboard report
Semi-Supervised Video Object Segmentation DAVIS 2016 PML F-measure (Mean) 79.3 #66 of 78 Archive leaderboard report
Semi-Supervised Video Object Segmentation DAVIS 2016 PML F-measure (Recall) 93.4 #66 of 78 Archive leaderboard report
Semi-Supervised Video Object Segmentation DAVIS 2016 PML J&F 77.4 #66 of 78 Archive leaderboard report
Semi-Supervised Video Object Segmentation DAVIS 2016 PML Jaccard (Decay) 8.5 #66 of 78 Archive leaderboard report
Semi-Supervised Video Object Segmentation DAVIS 2016 PML Jaccard (Mean) 75.5 #66 of 78 Archive leaderboard report
Semi-Supervised Video Object Segmentation DAVIS 2016 PML Jaccard (Recall) 89.6 #66 of 78 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections