{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/deepvo-towards-end-to-end-visual-odometry","title":"DeepVO: Towards End-to-End Visual Odometry with Deep Recurrent Convolutional Neural Networks","arxiv_id":"1709.08429","date":"2017-09-25","proceeding":null,"authors":["Sen Wang","Ronald Clark","Hongkai Wen","Niki Trigoni"],"abstract":"This paper studies monocular visual odometry (VO) problem. Most of existing\nVO algorithms are developed under a standard pipeline including feature\nextraction, feature matching, motion estimation, local optimisation, etc.\nAlthough some of them have demonstrated superior performance, they usually need\nto be carefully designed and specifically fine-tuned to work well in different\nenvironments. Some prior knowledge is also required to recover an absolute\nscale for monocular VO. This paper presents a novel end-to-end framework for\nmonocular VO by using deep Recurrent Convolutional Neural Networks (RCNNs).\nSince it is trained and deployed in an end-to-end manner, it infers poses\ndirectly from a sequence of raw RGB images (videos) without adopting any module\nin the conventional VO pipeline. Based on the RCNNs, it not only automatically\nlearns effective feature representation for the VO problem through\nConvolutional Neural Networks, but also implicitly models sequential dynamics\nand relations using deep Recurrent Neural Networks. Extensive experiments on\nthe KITTI VO dataset show competitive performance to state-of-the-art methods,\nverifying that the end-to-end Deep Learning technique can be a viable\ncomplement to the traditional VO systems.","url_abs":"http://arxiv.org/abs/1709.08429v1","url_pdf":"http://arxiv.org/pdf/1709.08429v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"deepvo-towards-end-to-end-visual-odometry","repo_url":"https://github.com/fshamshirdar/DeepVO","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"deepvo-towards-end-to-end-visual-odometry","repo_url":"https://github.com/jwolf02/rtdeepvo","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null},{"paper_slug":"deepvo-towards-end-to-end-visual-odometry","repo_url":"https://github.com/mrchanwoo/DeepVO-Keras","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null},{"paper_slug":"deepvo-towards-end-to-end-visual-odometry","repo_url":"https://github.com/ChiWeiHsiao/DeepVO-pytorch","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":null},{"paper_slug":"deepvo-towards-end-to-end-visual-odometry","repo_url":"https://github.com/remaro-network/Loss_VO_right","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"monocular-visual-odometry","task_name":"Monocular Visual Odometry"},{"task_slug":"motion-estimation","task_name":"Motion Estimation"},{"task_slug":"visual-odometry","task_name":"Visual Odometry"}],"methods":[{"method_slug":"1-bit-adam","method_name":"1-bit Adam"},{"method_slug":"adam","method_name":"Adam"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1709.08429","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}