{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/unsupervised-learning-of-depth-and-ego-motion-1","title":"Unsupervised Learning of Depth and Ego-Motion from Video","arxiv_id":"1704.07813","date":"2017-04-25","proceeding":"CVPR 2017 7","authors":["Tinghui Zhou","Matthew Brown","Noah Snavely","David G. Lowe"],"abstract":"We present an unsupervised learning framework for the task of monocular depth\nand camera motion estimation from unstructured video sequences. We achieve this\nby simultaneously training depth and camera pose estimation networks using the\ntask of view synthesis as the supervisory signal. The networks are thus coupled\nvia the view synthesis objective during training, but can be applied\nindependently at test time. Empirical evaluation on the KITTI dataset\ndemonstrates the effectiveness of our approach: 1) monocular depth performing\ncomparably with supervised methods that use either ground-truth pose or depth\nfor training, and 2) pose estimation performing favorably with established SLAM\nsystems under comparable input settings.","url_abs":"http://arxiv.org/abs/1704.07813v2","url_pdf":"http://arxiv.org/pdf/1704.07813v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"unsupervised-learning-of-depth-and-ego-motion-1","repo_url":"https://github.com/tinghuiz/SfMLearner","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"tf","reach":{"status":"unanswered"}},{"paper_slug":"unsupervised-learning-of-depth-and-ego-motion-1","repo_url":"https://github.com/ClementPinard/SfmLearner-Pytorch","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"camera-pose-estimation","task_name":"Camera Pose Estimation"},{"task_slug":"depth-and-camera-motion","task_name":"Depth And Camera Motion"},{"task_slug":"depth-estimation","task_name":"Depth Estimation"},{"task_slug":"motion-estimation","task_name":"Motion Estimation"},{"task_slug":"pose-estimation","task_name":"Pose Estimation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/camera-pose-estimation-on-kitti-odometry","task":"Camera Pose Estimation","dataset":"KITTI Odometry Benchmark","model":"SfMLearner","rank_in_archive_order":6,"of":7,"metrics":{"Absolute Trajectory Error [m]":"72.57","Average Rotational Error er[%]":"12.26","Average Translational Error et[%]":"29.78"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/1704.07813","atlas_url":"https://app.syntology.ai/?focus=1704.07813","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}