Papers › Self-Supervised Monocular Depth Estimation with Internal Feature Fusion

Self-Supervised Monocular Depth Estimation with Internal Feature Fusion

18 Oct 2021arXiv:2110.09482archive 2025-07-28

Hang Zhou, David Greenwood, Sarah Taylor

Self-supervised learning for depth estimation uses geometry in image sequences for supervision and shows promising results. Like many computer vision tasks, depth network performance is determined by the capability to learn accurate spatial and semantic representations from images. Therefore, it is natural to exploit semantic segmentation networks for depth estimation. In this work, based on a well-developed semantic segmentation network HRNet, we propose a novel depth estimation network DIFFNet, which can make use of semantic information in down and upsampling procedures. By applying feature fusion and an attention mechanism, our proposed method outperforms the state-of-the-art monocular depth estimation methods on the KITTI benchmark. Our method also demonstrates greater potential on higher resolution training data. We propose an additional extended evaluation strategy by establishing a test set of challenging cases, empirically derived from the standard benchmark.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

brandleyzhou/diffnet officialmentioned in papermentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Depth EstimationMonocular Depth EstimationSegmentationSelf-Supervised LearningUnsupervised Monocular Depth Estimation

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Monocular Depth Estimation KITTI Eigen split unsupervised DIFFNet (MS+1024x320) Delta < 1.25 0.911 #12 of 55 Archive leaderboard report
Monocular Depth Estimation KITTI Eigen split unsupervised DIFFNet (MS+1024x320) Delta < 1.25^2 0.968 #12 of 55 Archive leaderboard report
Monocular Depth Estimation KITTI Eigen split unsupervised DIFFNet (MS+1024x320) Delta < 1.25^3 0.984 #12 of 55 Archive leaderboard report
Monocular Depth Estimation KITTI Eigen split unsupervised DIFFNet (MS+1024x320) Mono X #12 of 55 Archive leaderboard report
Monocular Depth Estimation KITTI Eigen split unsupervised DIFFNet (MS+1024x320) RMSE 4.250 #12 of 55 Archive leaderboard report
Monocular Depth Estimation KITTI Eigen split unsupervised DIFFNet (MS+1024x320) RMSE log 0.172 #12 of 55 Archive leaderboard report
Monocular Depth Estimation KITTI Eigen split unsupervised DIFFNet (MS+1024x320) Sq Rel 0.678 #12 of 55 Archive leaderboard report
Monocular Depth Estimation KITTI Eigen split unsupervised DIFFNet (MS+1024x320) absolute relative error 0.094 #12 of 55 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Batch NormalizationConvolutionHRNetReLUResidual ConnectionTest

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections