Papers › H-Net: Unsupervised Attention-based Stereo Depth Estimation Leveraging Epipolar Geometry
H-Net: Unsupervised Attention-based Stereo Depth Estimation Leveraging Epipolar Geometry
Baoru Huang, Jian-Qing Zheng, Stamatia Giannarou, Daniel S. Elson
Depth estimation from a stereo image pair has become one of the most explored applications in computer vision, with most of the previous methods relying on fully supervised learning settings. However, due to the difficulty in acquiring accurate and scalable ground truth data, the training of fully supervised methods is challenging. As an alternative, self-supervised methods are becoming more popular to mitigate this challenge. In this paper, we introduce the H-Net, a deep-learning framework for unsupervised stereo depth estimation that leverages epipolar geometry to refine stereo matching. For the first time, a Siamese autoencoder architecture is used for depth estimation which allows mutual information between the rectified stereo images to be extracted. To enforce the epipolar constraint, the mutual epipolar attention mechanism has been designed which gives more emphasis to correspondences of features which lie on the same epipolar line while learning mutual information between the input stereo pair. Stereo correspondences are further enhanced by incorporating semantic information to the proposed attention mechanism. More specifically, the optimal transport algorithm is used to suppress attention and eliminate outliers in areas not visible in both cameras. Extensive experiments on KITTI2015 and Cityscapes show that our method outperforms the state-ofthe-art unsupervised stereo depth estimation methods while closing the gap with the fully supervised approaches.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
No code repository is listed for this paper in the archive or in Syntology's graph.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Depth Estimation | KITTI 2015 | H-Net (Ours) Full Eigen | Absolute relative error (AbsRel) | 0.076 | #1 of 2 | Archive leaderboard | report |
| Depth Estimation | KITTI 2015 | H-Net (Ours) Full Eigen | RMSE | 0.04025 | #1 of 2 | Archive leaderboard | report |
| Depth Estimation | KITTI 2015 | H-Net (Ours) Full Eigen | Sq Rel | 0.607 | #1 of 2 | Archive leaderboard | report |
| Depth Estimation | KITTI 2015 | H-Net (Ours) | Absolute relative error (AbsRel) | 0.094 | #2 of 2 | Archive leaderboard | report |
| Depth Estimation | KITTI 2015 | H-Net (Ours) | Sq Rel | 0.6 | #2 of 2 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections