{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/transformers-in-self-supervised-monocular","title":"Transformers in Self-Supervised Monocular Depth Estimation with Unknown Camera Intrinsics","arxiv_id":"2202.03131","date":"2022-02-07","proceeding":null,"authors":["Arnav Varma","Hemang Chawla","Bahram Zonooz","Elahe Arani"],"abstract":"The advent of autonomous driving and advanced driver assistance systems necessitates continuous developments in computer vision for 3D scene understanding. Self-supervised monocular depth estimation, a method for pixel-wise distance estimation of objects from a single camera without the use of ground truth labels, is an important task in 3D scene understanding. However, existing methods for this task are limited to convolutional neural network (CNN) architectures. In contrast with CNNs that use localized linear operations and lose feature resolution across the layers, vision transformers process at constant resolution with a global receptive field at every stage. While recent works have compared transformers against their CNN counterparts for tasks such as image classification, no study exists that investigates the impact of using transformers for self-supervised monocular depth estimation. Here, we first demonstrate how to adapt vision transformers for self-supervised monocular depth estimation. Thereafter, we compare the transformer and CNN-based architectures for their performance on KITTI depth prediction benchmarks, as well as their robustness to natural corruptions and adversarial attacks, including when the camera intrinsics are unknown. Our study demonstrates how transformer-based architecture, though lower in run-time efficiency, achieves comparable performance while being more robust and generalizable.","url_abs":"https://arxiv.org/abs/2202.03131v1","url_pdf":"https://arxiv.org/pdf/2202.03131v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"transformers-in-self-supervised-monocular","repo_url":"https://github.com/neurai-lab/mt-sfmlearner","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"autonomous-driving","task_name":"Autonomous Driving"},{"task_slug":"depth-estimation","task_name":"Depth Estimation"},{"task_slug":"depth-prediction","task_name":"Depth Prediction"},{"task_slug":"image-classification","task_name":"Image Classification"},{"task_slug":"monocular-depth-estimation","task_name":"Monocular Depth Estimation"},{"task_slug":"scene-understanding","task_name":"Scene Understanding"},{"task_slug":"image-classification","task_name":"image-classification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2202.03131","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2202.03131"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/neurai-lab/mt-sfmlearner","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"unverified":6},"by_repo_kind":{"listed":{"samples":6,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"275e5e769187fecc","entry":"SSIM","repo":"neurai-lab/mt-sfmlearner","repo_kind":"listed","path":"mtsfmlearner/losses/multiview_photometric_loss.py","file_url":"https://github.com/neurai-lab/mt-sfmlearner/blob/HEAD/mtsfmlearner/losses/multiview_photometric_loss.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"275e5e769187fecc"}},{"code_sha256_prefix":"1c9d1d852e11deaf","entry":"read_npz_depth","repo":"neurai-lab/mt-sfmlearner","repo_kind":"listed","path":"mtsfmlearner/datasets/kitti_dataset.py","file_url":"https://github.com/neurai-lab/mt-sfmlearner/blob/HEAD/mtsfmlearner/datasets/kitti_dataset.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"1c9d1d852e11deaf"}},{"code_sha256_prefix":"837c6a474dd26eff","entry":"resize_depth","repo":"neurai-lab/mt-sfmlearner","repo_kind":"listed","path":"mtsfmlearner/datasets/augmentations.py","file_url":"https://github.com/neurai-lab/mt-sfmlearner/blob/HEAD/mtsfmlearner/datasets/augmentations.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"837c6a474dd26eff"}},{"code_sha256_prefix":"5c94a7e309621404","entry":"rotx","repo":"neurai-lab/mt-sfmlearner","repo_kind":"listed","path":"mtsfmlearner/datasets/kitti_dataset_utils.py","file_url":"https://github.com/neurai-lab/mt-sfmlearner/blob/HEAD/mtsfmlearner/datasets/kitti_dataset_utils.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"5c94a7e309621404"}},{"code_sha256_prefix":"0a19721bdaffaddd","entry":"roty","repo":"neurai-lab/mt-sfmlearner","repo_kind":"listed","path":"mtsfmlearner/datasets/kitti_dataset_utils.py","file_url":"https://github.com/neurai-lab/mt-sfmlearner/blob/HEAD/mtsfmlearner/datasets/kitti_dataset_utils.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"0a19721bdaffaddd"}},{"code_sha256_prefix":"e1451f5589db8939","entry":"rotz","repo":"neurai-lab/mt-sfmlearner","repo_kind":"listed","path":"mtsfmlearner/datasets/kitti_dataset_utils.py","file_url":"https://github.com/neurai-lab/mt-sfmlearner/blob/HEAD/mtsfmlearner/datasets/kitti_dataset_utils.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"e1451f5589db8939"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}