{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/mvsnet-depth-inference-for-unstructured-multi","title":"MVSNet: Depth Inference for Unstructured Multi-view Stereo","arxiv_id":"1804.02505","date":"2018-04-07","proceeding":"ECCV 2018 9","authors":["Yao Yao","Zixin Luo","Shiwei Li","Tian Fang","Long Quan"],"abstract":"We present an end-to-end deep learning architecture for depth map inference\nfrom multi-view images. In the network, we first extract deep visual image\nfeatures, and then build the 3D cost volume upon the reference camera frustum\nvia the differentiable homography warping. Next, we apply 3D convolutions to\nregularize and regress the initial depth map, which is then refined with the\nreference image to generate the final output. Our framework flexibly adapts\narbitrary N-view inputs using a variance-based cost metric that maps multiple\nfeatures into one cost feature. The proposed MVSNet is demonstrated on the\nlarge-scale indoor DTU dataset. With simple post-processing, our method not\nonly significantly outperforms previous state-of-the-arts, but also is several\ntimes faster in runtime. We also evaluate MVSNet on the complex outdoor Tanks\nand Temples dataset, where our method ranks first before April 18, 2018 without\nany fine-tuning, showing the strong generalization ability of MVSNet.","url_abs":"http://arxiv.org/abs/1804.02505v2","url_pdf":"http://arxiv.org/pdf/1804.02505v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"mvsnet-depth-inference-for-unstructured-multi","repo_url":"https://github.com/Skoltech-3D/sk3d_data","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"NOASSERTION"}},{"paper_slug":"mvsnet-depth-inference-for-unstructured-multi","repo_url":"https://github.com/YoYo000/MVSNet","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"mvsnet-depth-inference-for-unstructured-multi","repo_url":"https://github.com/kwea123/MVSNet_pl","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"mvsnet-depth-inference-for-unstructured-multi","repo_url":"https://github.com/xy-guo/MVSNet_pytorch","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"mvsnet-depth-inference-for-unstructured-multi","repo_url":"https://github.com/2023-MindSpore-1/ms-code-213/tree/main/ESRGAN","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null}],"tasks":[{"task_slug":"3d-reconstruction","task_name":"3D Reconstruction"},{"task_slug":"point-clouds","task_name":"Point Clouds"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/3d-reconstruction-on-dtu","task":"3D Reconstruction","dataset":"DTU","model":"MVSNet","rank_in_archive_order":21,"of":24,"metrics":{"Acc":"0.396","Comp":"0.527","Overall":"0.462"},"uses_additional_data":false},{"leaderboard":"/sota/point-clouds-on-tanks-and-temples","task":"Point Clouds","dataset":"Tanks and Temples","model":"MVSNet","rank_in_archive_order":21,"of":21,"metrics":{"Mean F1 (Intermediate)":"43.48"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1804.02505","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}