{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/salientdso-bringing-attention-to-direct","title":"SalientDSO: Bringing Attention to Direct Sparse Odometry","arxiv_id":"1803.00127","date":"2018-02-28","proceeding":null,"authors":["Huai-Jen Liang","Nitin J. Sanket","Cornelia Fermüller","Yiannis Aloimonos"],"abstract":"Although cluttered indoor scenes have a lot of useful high-level semantic\ninformation which can be used for mapping and localization, most Visual\nOdometry (VO) algorithms rely on the usage of geometric features such as\npoints, lines and planes. Lately, driven by this idea, the joint optimization\nof semantic labels and obtaining odometry has gained popularity in the robotics\ncommunity. The joint optimization is good for accurate results but is generally\nvery slow. At the same time, in the vision community, direct and sparse\napproaches for VO have stricken the right balance between speed and accuracy.\n  We merge the successes of these two communities and present a way to\nincorporate semantic information in the form of visual saliency to Direct\nSparse Odometry - a highly successful direct sparse VO algorithm. We also\npresent a framework to filter the visual saliency based on scene parsing. Our\nframework, SalientDSO, relies on the widely successful deep learning based\napproaches for visual saliency and scene parsing which drives the feature\nselection for obtaining highly-accurate and robust VO even in the presence of\nas few as 40 point features per frame. We provide extensive quantitative\nevaluation of SalientDSO on the ICL-NUIM and TUM monoVO datasets and show that\nwe outperform DSO and ORB-SLAM - two very popular state-of-the-art approaches\nin the literature. We also collect and publicly release a CVL-UMD dataset which\ncontains two indoor cluttered sequences on which we show qualitative\nevaluations. To our knowledge this is the first paper to use visual saliency\nand scene parsing to drive the feature selection in direct VO.","url_abs":"http://arxiv.org/abs/1803.00127v1","url_pdf":"http://arxiv.org/pdf/1803.00127v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"salientdso-bringing-attention-to-direct","repo_url":"https://github.com/prgumd/SalientDSO","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"tf","reach":null}],"tasks":[{"task_slug":"scene-parsing","task_name":"Scene Parsing"},{"task_slug":"visual-odometry","task_name":"Visual Odometry"},{"task_slug":"feature-selection","task_name":"feature selection"}],"methods":[{"method_slug":"speed","method_name":"SPEED"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}