{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/every-pixel-counts-unsupervised-geometry","title":"Every Pixel Counts: Unsupervised Geometry Learning with Holistic 3D Motion Understanding","arxiv_id":"1806.10556","date":"2018-06-27","proceeding":null,"authors":["Zhenheng Yang","Peng Wang","Yang Wang","Wei Xu","Ram Nevatia"],"abstract":"Learning to estimate 3D geometry in a single image by watching unlabeled\nvideos via deep convolutional network has made significant process recently.\nCurrent state-of-the-art (SOTA) methods, are based on the learning framework of\nrigid structure-from-motion, where only 3D camera ego motion is modeled for\ngeometry estimation.However, moving objects also exist in many videos, e.g.\nmoving cars in a street scene. In this paper, we tackle such motion by\nadditionally incorporating per-pixel 3D object motion into the learning\nframework, which provides holistic 3D scene flow understanding and helps single\nimage geometry estimation. Specifically, given two consecutive frames from a\nvideo, we adopt a motion network to predict their relative 3D camera pose and a\nsegmentation mask distinguishing moving objects and rigid background. An\noptical flow network is used to estimate dense 2D per-pixel correspondence. A\nsingle image depth network predicts depth maps for both images. The four types\nof information, i.e. 2D flow, camera pose, segment mask and depth maps, are\nintegrated into a differentiable holistic 3D motion parser (HMP), where\nper-pixel 3D motion for rigid background and moving objects are recovered. We\ndesign various losses w.r.t. the two types of 3D motions for training the depth\nand motion networks, yielding further error reduction for estimated geometry.\nFinally, in order to solve the 3D motion confusion from monocular videos, we\ncombine stereo images into joint training. Experiments on KITTI 2015 dataset\nshow that our estimated geometry, 3D motion and moving object masks, not only\nare constrained to be consistent, but also significantly outperforms other SOTA\nalgorithms, demonstrating the benefits of our approach.","url_abs":"http://arxiv.org/abs/1806.10556v2","url_pdf":"http://arxiv.org/pdf/1806.10556v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"3d-geometry","task_name":"3D geometry"},{"task_slug":"depth-and-camera-motion","task_name":"Depth And Camera Motion"},{"task_slug":"depth-estimation","task_name":"Depth Estimation"},{"task_slug":"optical-flow-estimation","task_name":"Optical Flow Estimation"},{"task_slug":"scene-flow-estimation","task_name":"Scene Flow Estimation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/scene-flow-estimation-on-kitti-2015-scene","task":"Scene Flow Estimation","dataset":"KITTI 2015 Scene Flow Training","model":"EPC","rank_in_archive_order":2,"of":4,"metrics":{" Runtime (s)":"0.05","D1-all":"26.81","D2-all":"60.97","Fl-all":"25.74","SF-all":"(>60.97)"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1806.10556","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}