Papers › Self-supervised Learning for Video Correspondence Flow

Self-supervised Learning for Video Correspondence Flow

2 May 2019arXiv:1905.00875archive 2025-07-28

Zihang Lai, Weidi Xie

The objective of this paper is self-supervised learning of feature embeddings that are suitable for matching correspondences along the videos, which we term correspondence flow. By leveraging the natural spatial-temporal coherence in videos, we propose to train a ``pointer'' that reconstructs a target frame by copying pixels from a reference frame. We make the following contributions: First, we introduce a simple information bottleneck that forces the model to learn robust features for correspondence matching, and prevent it from learning trivial solutions, \eg matching based on low-level colour information. Second, to tackle the challenges from tracker drifting, due to complex object deformations, illumination changes and occlusions, we propose to train a recursive model over long temporal windows with scheduled sampling and cycle consistency. Third, we achieve state-of-the-art performance on DAVIS 2017 video segmentation and JHMDB keypoint tracking tasks, outperforming all previous self-supervised learning approaches by a significant margin. Fourth, in order to shed light on the potential of self-supervised learning on the task of video correspondence flow, we probe the upper bound by training on additional data, \ie more diverse videos, further demonstrating significant improvements on video segmentation.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

zlai0/CorrFlow officialmentioned in papermentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Self-Supervised LearningSemi-Supervised Video Object SegmentationUnsupervised Video Object SegmentationVideo Correspondence FlowVideo SegmentationVideo Semantic Segmentation

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Semi-Supervised Video Object Segmentation DAVIS 2017 (val) CorrFlow F-measure (Mean) 52.2 #79 of 81 Archive leaderboard report
Semi-Supervised Video Object Segmentation DAVIS 2017 (val) CorrFlow F-measure (Recall) 56.0 #79 of 81 Archive leaderboard report
Semi-Supervised Video Object Segmentation DAVIS 2017 (val) CorrFlow J&F 50.3 #79 of 81 Archive leaderboard report
Semi-Supervised Video Object Segmentation DAVIS 2017 (val) CorrFlow Jaccard (Mean) 48.4 #79 of 81 Archive leaderboard report
Semi-Supervised Video Object Segmentation DAVIS 2017 (val) CorrFlow Jaccard (Recall) 53.2 #79 of 81 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections