{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/transdssl-transformer-based-depth-estimation","title":"TransDSSL: Transformer based Depth Estimation via Self-Supervised Learning","arxiv_id":null,"date":"2022-08-05","proceeding":"journal 2022 8","authors":["Daechan Han","Jeongmin Shin","Namil Kim","Soomnim Hwang","Yukyung Choi"],"abstract":"Recently, transformers have been widely adopted for various computer vision tasks and show promising results due to their ability to encode long-range spatial dependencies in an image effectively. However, very few studies on adopting transformers in self-supervised depth estimation have been conducted. When replacing the CNN architecture with the transformer in self-supervised learning of depth, we encounter several problems such as problematic multi-scale photometric loss function when used with transformers and, insuffcient ability to capture local details. In this paper, we propose an attention-based decoder module, Pixel-Wise Skip Attention (PWSA), to enhance fine details in feature maps while keeping global context from transformers. In addition, we propose utilizing self-distillation loss with single-scale photometric loss to alleviate the instability of transformer training by using correct training signals. We demonstrate that the proposed model performs accurate predictions on large objects and thin structures that require global context and local details. Our model achieves state-ofthe-art performance among the self-supervised monocular depth estimation methods on KITTI and DDAD benchmarks","url_abs":"https://ieeexplore.ieee.org/document/9851497","url_pdf":"https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=9851497","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"transdssl-transformer-based-depth-estimation","repo_url":"https://github.com/sejong-rcv/2021.Paper.TransDSSL","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"decoder","task_name":"Decoder"},{"task_slug":"depth-estimation","task_name":"Depth Estimation"},{"task_slug":"monocular-depth-estimation","task_name":"Monocular Depth Estimation"},{"task_slug":"self-supervised-learning","task_name":"Self-Supervised Learning"},{"task_slug":"unsupervised-monocular-depth-estimation","task_name":"Unsupervised Monocular Depth Estimation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/monocular-depth-estimation-on-ddad","task":"Monocular Depth Estimation","dataset":"DDAD","model":"TransDSSL","rank_in_archive_order":4,"of":4,"metrics":{"RMSE":"14.350","RMSE log":"0.172","Sq Rel":"3.591","absolute relative error":"0.151"},"uses_additional_data":false},{"leaderboard":"/sota/monocular-depth-estimation-on-kitti-eigen-1","task":"Monocular Depth Estimation","dataset":"KITTI Eigen split unsupervised","model":"TransDSSL","rank_in_archive_order":14,"of":55,"metrics":{"Delta < 1.25":"0.906","Delta < 1.25^2":"0.967","Delta < 1.25^3":"0.984","Mono":"O","RMSE":"4.321","RMSE log":"0.172","Sq Rel":"0.711","absolute relative error":"0.095"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}