Papers › MeSa: Masked, Geometric, and Supervised Pre-training for Monocular Depth Estimation

MeSa: Masked, Geometric, and Supervised Pre-training for Monocular Depth Estimation

6 Oct 2023arXiv:2310.04551archive 2025-07-28

Muhammad Osama Khan, Junbang Liang, Chun-Kai Wang, Shan Yang, Yu Lou

Pre-training has been an important ingredient in developing strong monocular depth estimation models in recent years. For instance, self-supervised learning (SSL) is particularly effective by alleviating the need for large datasets with dense ground-truth depth maps. However, despite these improvements, our study reveals that the later layers of the SOTA SSL method are actually suboptimal. By examining the layer-wise representations, we demonstrate significant changes in these later layers during fine-tuning, indicating the ineffectiveness of their pre-trained features for depth estimation. To address these limitations, we propose MeSa, a comprehensive framework that leverages the complementary strengths of masked, geometric, and supervised pre-training. Hence, MeSa benefits from not only general-purpose representations learnt via masked pre training but also specialized depth-specific features acquired via geometric and supervised pre-training. Our CKA layer-wise analysis confirms that our pre-training strategy indeed produces improved representations for the later layers, overcoming the drawbacks of the SOTA SSL method. Furthermore, via experiments on the NYUv2 and IBims-1 datasets, we demonstrate that these enhanced representations translate to performance improvements in both the in-distribution and out-of-distribution settings. We also investigate the influence of the pre-training dataset and demonstrate the efficacy of pre-training on LSUN, which yields significantly better pre-trained representations. Overall, our approach surpasses the masked pre-training SSL method by a substantial margin of 17.1% on the RMSE. Moreover, even without utilizing any recently proposed techniques, MeSa also outperforms the most recent methods and establishes a new state-of-the-art for monocular depth estimation on the challenging NYUv2 dataset.

PaperPDF

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Depth EstimationMonocular Depth EstimationSelf-Supervised Learning

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Monocular Depth Estimation NYU-Depth V2 MeSa Delta < 1.25 0.964 #19 of 85 Archive leaderboard report
Monocular Depth Estimation NYU-Depth V2 MeSa Delta < 1.25^2 0.995 #19 of 85 Archive leaderboard report
Monocular Depth Estimation NYU-Depth V2 MeSa Delta < 1.25^3 0.999 #19 of 85 Archive leaderboard report
Monocular Depth Estimation NYU-Depth V2 MeSa RMSE 0.238 #19 of 85 Archive leaderboard report
Monocular Depth Estimation NYU-Depth V2 MeSa absolute relative error 0.066 #19 of 85 Archive leaderboard report
Monocular Depth Estimation NYU-Depth V2 MeSa log 10 0.029 #19 of 85 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections