Papers › NVDS+: Towards Efficient and Versatile Neural Stabilizer for Video Depth Estimation

NVDS+: Towards Efficient and Versatile Neural Stabilizer for Video Depth Estimation

17 Jul 2023ICCV 2023 1arXiv:2307.08695archive 2025-07-28

Yiran Wang, Min Shi, Jiaqi Li, Chaoyi Hong, Zihao Huang, Juewen Peng, Zhiguo Cao, Jianming Zhang, Ke Xian, Guosheng Lin

Video depth estimation aims to infer temporally consistent depth. One approach is to finetune a single-image model on each video with geometry constraints, which proves inefficient and lacks robustness. An alternative is learning to enforce consistency from data, which requires well-designed models and sufficient video depth data. To address both challenges, we introduce NVDS+ that stabilizes inconsistent depth estimated by various single-image models in a plug-and-play manner. We also elaborate a large-scale Video Depth in the Wild (VDW) dataset, which contains 14,203 videos with over two million frames, making it the largest natural-scene video depth dataset. Additionally, a bidirectional inference strategy is designed to improve consistency by adaptively fusing forward and backward predictions. We instantiate a model family ranging from small to large scales for different applications. The method is evaluated on VDW dataset and three public benchmarks. To further prove the versatility, we extend NVDS+ to video semantic segmentation and several downstream applications like bokeh rendering, novel view synthesis, and 3D reconstruction. Experimental results show that our method achieves significant improvements in consistency, accuracy, and efficiency. Our work serves as a solid baseline and data foundation for learning-based video depth estimation. Code and dataset are available at: https://github.com/RaymondWang987/NVDS

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

raymondwang987/nvds officialmentioned in papermentioned on GitHubpytorchMIT report
raymondwang987/fmnet mentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

3D ReconstructionDepth EstimationMonocular Depth EstimationNovel View SynthesisSemantic SegmentationVideo Semantic Segmentation

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Monocular Depth Estimation NYU-Depth V2 NVDS(DPT-L) Delta < 1.25 0.9493 #22 of 85 Archive leaderboard report
Monocular Depth Estimation NYU-Depth V2 NVDS(DPT-L) Delta < 1.25^2 0.991 #22 of 85 Archive leaderboard report
Monocular Depth Estimation NYU-Depth V2 NVDS(DPT-L) Delta < 1.25^3 0.997 #22 of 85 Archive leaderboard report
Monocular Depth Estimation NYU-Depth V2 NVDS(DPT-L) RMSE 0.282 #22 of 85 Archive leaderboard report
Monocular Depth Estimation NYU-Depth V2 NVDS(DPT-L) absolute relative error 0.072 #22 of 85 Archive leaderboard report
Monocular Depth Estimation NYU-Depth V2 NVDS(DPT-L) log 10 0.031 #22 of 85 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections