{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/beyond-photometric-loss-for-self-supervised","title":"Beyond Photometric Loss for Self-Supervised Ego-Motion Estimation","arxiv_id":"1902.09103","date":"2019-02-25","proceeding":null,"authors":["Tianwei Shen","Zixin Luo","Lei Zhou","Hanyu Deng","Runze Zhang","Tian Fang","Long Quan"],"abstract":"Accurate relative pose is one of the key components in visual odometry (VO)\nand simultaneous localization and mapping (SLAM). Recently, the self-supervised\nlearning framework that jointly optimizes the relative pose and target image\ndepth has attracted the attention of the community. Previous works rely on the\nphotometric error generated from depths and poses between adjacent frames,\nwhich contains large systematic error under realistic scenes due to reflective\nsurfaces and occlusions. In this paper, we bridge the gap between geometric\nloss and photometric loss by introducing the matching loss constrained by\nepipolar geometry in a self-supervised framework. Evaluated on the KITTI\ndataset, our method outperforms the state-of-the-art unsupervised ego-motion\nestimation methods by a large margin. The code and data are available at\nhttps://github.com/hlzz/DeepMatchVO.","url_abs":"http://arxiv.org/abs/1902.09103v1","url_pdf":"http://arxiv.org/pdf/1902.09103v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"beyond-photometric-loss-for-self-supervised","repo_url":"https://github.com/hlzz/DeepMatchVO","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"camera-pose-estimation","task_name":"Camera Pose Estimation"},{"task_slug":"motion-estimation","task_name":"Motion Estimation"},{"task_slug":"self-supervised-learning","task_name":"Self-Supervised Learning"},{"task_slug":"simultaneous-localization-and-mapping","task_name":"Simultaneous Localization and Mapping"},{"task_slug":"visual-odometry","task_name":"Visual Odometry"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/camera-pose-estimation-on-kitti-odometry","task":"Camera Pose Estimation","dataset":"KITTI Odometry Benchmark","model":"DeepMatchVO","rank_in_archive_order":3,"of":7,"metrics":{"Absolute Trajectory Error [m]":"25.76","Average Rotational Error er[%]":"4.85","Average Translational Error et[%]":"11.05"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1902.09103","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}