{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/zoom-and-learn-generalizing-deep-stereo","title":"Zoom and Learn: Generalizing Deep Stereo Matching to Novel Domains","arxiv_id":"1803.06641","date":"2018-03-18","proceeding":"CVPR 2018 6","authors":["Jiahao Pang","Wenxiu Sun","Chengxi Yang","Jimmy Ren","Ruichao Xiao","Jin Zeng","Liang Lin"],"abstract":"Despite the recent success of stereo matching with convolutional neural\nnetworks (CNNs), it remains arduous to generalize a pre-trained deep stereo\nmodel to a novel domain. A major difficulty is to collect accurate ground-truth\ndisparities for stereo pairs in the target domain. In this work, we propose a\nself-adaptation approach for CNN training, utilizing both synthetic training\ndata (with ground-truth disparities) and stereo pairs in the new domain\n(without ground-truths). Our method is driven by two empirical observations. By\nfeeding real stereo pairs of different domains to stereo models pre-trained\nwith synthetic data, we see that: i) a pre-trained model does not generalize\nwell to the new domain, producing artifacts at boundaries and ill-posed\nregions; however, ii) feeding an up-sampled stereo pair leads to a disparity\nmap with extra details. To avoid i) while exploiting ii), we formulate an\niterative optimization problem with graph Laplacian regularization. At each\niteration, the CNN adapts itself better to the new domain: we let the CNN learn\nits own higher-resolution output; at the meanwhile, a graph Laplacian\nregularization is imposed to discriminatively keep the desired edges while\nsmoothing out the artifacts. We demonstrate the effectiveness of our method in\ntwo domains: daily scenes collected by smartphone cameras, and street views\ncaptured in a driving car.","url_abs":"http://arxiv.org/abs/1803.06641v1","url_pdf":"http://arxiv.org/pdf/1803.06641v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"zoom-and-learn-generalizing-deep-stereo","repo_url":"https://github.com/Artifineuro/zole","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"none","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":"stereo-matching-1","task_name":"Stereo Matching"},{"task_slug":"stereo-matching","task_name":"Stereo Matching Hand"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1803.06641","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}