Papers › CroCo v2: Improved Cross-view Completion Pre-training for Stereo Matching and Optical Flow

CroCo v2: Improved Cross-view Completion Pre-training for Stereo Matching and Optical Flow

18 Nov 2022ICCV 2023 1arXiv:2211.10408archive 2025-07-28

Philippe Weinzaepfel, Thomas Lucas, Vincent Leroy, Yohann Cabon, Vaibhav Arora, Romain Brégier, Gabriela Csurka, Leonid Antsfeld, Boris Chidlovskii, Jérôme Revaud

Despite impressive performance for high-level downstream tasks, self-supervised pre-training methods have not yet fully delivered on dense geometric vision tasks such as stereo matching or optical flow. The application of self-supervised concepts, such as instance discrimination or masked image modeling, to geometric tasks is an active area of research. In this work, we build on the recent cross-view completion framework, a variation of masked image modeling that leverages a second view from the same scene which makes it well suited for binocular downstream tasks. The applicability of this concept has so far been limited in at least two ways: (a) by the difficulty of collecting real-world image pairs -- in practice only synthetic data have been used -- and (b) by the lack of generalization of vanilla transformers to dense downstream tasks for which relative position is more meaningful than absolute position. We explore three avenues of improvement. First, we introduce a method to collect suitable real-world image pairs at large scale. Second, we experiment with relative positional embeddings and show that they enable vision transformers to perform substantially better. Third, we scale up vision transformer based cross-completion architectures, which is made possible by the use of large amounts of data. With these improvements, we show for the first time that state-of-the-art results on stereo matching and optical flow can be reached without using any classical task-specific techniques like correlation volume, iterative estimation, image warping or multi-scale reasoning, thus paving the way towards universal vision models.

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

naver/croco officialmentioned in papermentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Optical Flow EstimationSelf-Supervised LearningStereo Matching

1 archive task tag without a task page not shown.

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Optical Flow Estimation KITTI 2012 CroCo-Flow Average End-Point Error 0.8 #1 of 12 Archive leaderboard report
Optical Flow Estimation KITTI 2012 CroCo-Flow Noc 0.5 #1 of 12 Archive leaderboard report
Optical Flow Estimation KITTI 2012 CroCo-Flow Out-Noc 1.57 #1 of 12 Archive leaderboard report
Optical Flow Estimation KITTI 2015 CroCo-Flow Fl-all 3.64 #4 of 18 Archive leaderboard report
Optical Flow Estimation KITTI 2015 CroCo-Flow Fl-fg 5.94 #4 of 18 Archive leaderboard report
Optical Flow Estimation Sintel-clean CroCo-Flow Average End-Point Error 1.092 #4 of 29 Archive leaderboard report
Optical Flow Estimation Sintel-final CroCo-Flow Average End-Point Error 2.436 #5 of 28 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

AttentionDense ConnectionsLayer NormalizationLinear LayerMulti-Head AttentionResidual ConnectionSoftmaxVision Transformer

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections