{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/web-stereo-video-supervision-for-depth","title":"Web Stereo Video Supervision for Depth Prediction from Dynamic Scenes","arxiv_id":"1904.11112","date":"2019-04-25","proceeding":null,"authors":["Chaoyang Wang","Simon Lucey","Federico Perazzi","Oliver Wang"],"abstract":"We present a fully data-driven method to compute depth from diverse monocular\nvideo sequences that contain large amounts of non-rigid objects, e.g., people.\nIn order to learn reconstruction cues for non-rigid scenes, we introduce a new\ndataset consisting of stereo videos scraped in-the-wild. This dataset has a\nwide variety of scene types, and features large amounts of nonrigid objects,\nespecially people. From this, we compute disparity maps to be used as\nsupervision to train our approach. We propose a loss function that allows us to\ngenerate a depth prediction even with unknown camera intrinsics and stereo\nbaselines in the dataset. We validate the use of large amounts of Internet\nvideo by evaluating our method on existing video datasets with depth\nsupervision, including SINTEL, and KITTI, and show that our approach\ngeneralizes better to natural scenes.","url_abs":"http://arxiv.org/abs/1904.11112v1","url_pdf":"http://arxiv.org/pdf/1904.11112v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"depth-estimation","task_name":"Depth Estimation"},{"task_slug":"depth-prediction","task_name":"Depth Prediction"}],"methods":[],"datasets_introduced":[{"slug":"wsvd","name":"WSVD","full_name":"Web Stereo Video Dataset"}],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1904.11112","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}