Papers › Attention Attention Everywhere: Monocular Depth Prediction with Skip Attention

Attention Attention Everywhere: Monocular Depth Prediction with Skip Attention

17 Oct 2022arXiv:2210.09071archive 2025-07-28

Ashutosh Agarwal, Chetan Arora

Monocular Depth Estimation (MDE) aims to predict pixel-wise depth given a single RGB image. For both, the convolutional as well as the recent attention-based models, encoder-decoder-based architectures have been found to be useful due to the simultaneous requirement of global context and pixel-level resolution. Typically, a skip connection module is used to fuse the encoder and decoder features, which comprises of feature map concatenation followed by a convolution operation. Inspired by the demonstrated benefits of attention in a multitude of computer vision problems, we propose an attention-based fusion of encoder and decoder features. We pose MDE as a pixel query refinement problem, where coarsest-level encoder features are used to initialize pixel-level queries, which are then refined to higher resolutions by the proposed Skip Attention Module (SAM). We formulate the prediction problem as ordinal regression over the bin centers that discretize the continuous depth range and introduce a Bin Center Predictor (BCP) module that predicts bins at the coarsest level using pixel queries. Apart from the benefit of image adaptive depth binning, the proposed design helps learn improved depth embedding in initial pixel queries via direct supervision from the ground truth. Extensive experiments on the two canonical datasets, NYUV2 and KITTI, show that our architecture outperforms the state-of-the-art by 5.3% and 3.9%, respectively, along with an improved generalization performance by 9.4% on the SUNRGBD dataset. Code is available at https://github.com/ashutosh1807/PixelFormer.git.

PaperPDFCodeCode Syntology ran

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

For agents, Syntology's MCP tool lists every function and class Syntology harvested from this paper and whether it ran (how to connect): get_harvested_code_for_paper(arxiv_id="2210.09071")

Code

Syntology Ran 5 of 11 code samples harvested from 1 repository linked to this paper; 6 have no recorded run. Of those that ran: 1 ran · honoured contract; 1 ran · our draft was wrong; 1 ran · fixture could not drive it; 2 ran with no contract checked.

By repository: official repository: 11 samples from 1 repository, 5 ran. The run record, sample by sample. “Ran” means executed on a synthesized input, not that the code is correct or reproduces the paper.

ashutosh1807/pixelformer officialmentioned in papermentioned on GitHubpytorchMIT report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

11 samples harvested; 5 ran; 1 honoured the contract we drafted; 6 have no recorded run. Read from Syntology's graph 2026-09-24; that is when this build read the record, not when the samples ran.

1ran · honoured contract
1ran · our draft was wrong
1ran · fixture could not drive it
2ran
6unverified

Licence: 11 of the 11 samples are pointer only, meaning Syntology does not serve that copy's text. This page shows no code text for any sample; each one links to its file in the repository.

Harvested from ashutosh1807/pixelformer. “Ran” means the sample executed on a synthesized input. It does not mean the output is correct, and nothing here reproduces the paper's results. “Honoured” and “violated” refer to a contract Syntology drafted from the code itself; “our draft was wrong” and “fixture could not drive it” are failures of Syntology's instrument, not of the code.

Each sample ends with its code_sha256, Syntology's identity for that exact code. An agent fetches the stored sample with Syntology's MCP tool get_code(code_sha256="…") (how to connect); click an identity to copy that call.

Repository labels, per sample. official repository: The archive marks this repository official for the paper. named in the paper: The archive records that the paper mentions this repository; it is not marked official. community (archive-listed): In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper. found in paper text by Syntology: Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted. community: Not in the archive's code links for this paper; a community repository Syntology harvested. Samples from a repository marked official are listed first. Licence labels name the repository's licence as recorded at harvest. “Pointer only” means Syntology does not serve that copy's text, for one of four reasons: no licence file was found; the licence was not identified; the licence is recorded as permissive but that copy's record is not marked cleared; or the licence is outside the permissive list Syntology serves text under (MIT, Apache-2.0, BSD and similar). Some licences outside that list permit redistribution, such as WTFPL, and GPL-3.0 under its conditions; they are simply not on the list. Hover a licence label for the reason. File links open the file on GitHub at the default branch, which may have changed since the harvest.

window_partition ashutosh1807/pixelformer/pixelformer/networks/SAM.py official repository ran · fixture could not drive it fingerprinted MIT recorded; this copy not marked cleared · pointer only · 144d10b49baeb8a6 · report
get_num_lines ashutosh1807/pixelformer/pixelformer/utils.py official repository ran · honoured contract MIT recorded; this copy not marked cleared · pointer only · 050c56fb7282eb88 · report
normalize_result ashutosh1807/pixelformer/pixelformer/utils.py official repository ran MIT recorded; this copy not marked cleared · pointer only · 8a05ea4d5e917bf4 · report
upsample ashutosh1807/pixelformer/pixelformer/networks/PixelFormer.py official repository ran MIT recorded; this copy not marked cleared · pointer only · 35e6bae6a3fe0669 · report
window_reverse ashutosh1807/pixelformer/pixelformer/networks/SAM.py official repository ran · our draft was wrong MIT recorded; this copy not marked cleared · pointer only · 61bf152e6a42a184 · report
colorize ashutosh1807/pixelformer/pixelformer/utils.py official repository unverified MIT recorded; this copy not marked cleared · pointer only · 9999af2cf9c1b6bb · report
is_module_wrapper ashutosh1807/pixelformer/pixelformer/networks/utils.py official repository unverified MIT recorded; this copy not marked cleared · pointer only · 9b9dd29ad0d626a0 · report
load_url_dist ashutosh1807/pixelformer/pixelformer/networks/utils.py official repository unverified MIT recorded; this copy not marked cleared · pointer only · 92dad73ee855852f · report
preprocessing_transforms ashutosh1807/pixelformer/pixelformer/dataloaders/dataloader_kittipred.py official repository unverified MIT recorded; this copy not marked cleared · pointer only · ea623328271af85a · report
preprocessing_transforms ashutosh1807/pixelformer/pixelformer/dataloaders/dataloader.py official repository unverified MIT recorded; this copy not marked cleared · pointer only · 5041fa3a9230ea8a · report
resize ashutosh1807/pixelformer/pixelformer/networks/utils.py official repository unverified MIT recorded; this copy not marked cleared · pointer only · 9012d3fdd4cbaea1 · report

Tasks

DecoderDepth EstimationDepth PredictionMonocular Depth EstimationPrediction

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Monocular Depth Estimation KITTI Eigen split PixelFormer Delta < 1.25 0.976 #24 of 79 Archive leaderboard report
Monocular Depth Estimation KITTI Eigen split PixelFormer Delta < 1.25^2 0.997 #24 of 79 Archive leaderboard report
Monocular Depth Estimation KITTI Eigen split PixelFormer Delta < 1.25^3 0.999 #24 of 79 Archive leaderboard report
Monocular Depth Estimation KITTI Eigen split PixelFormer RMSE 2.081 #24 of 79 Archive leaderboard report
Monocular Depth Estimation KITTI Eigen split PixelFormer RMSE log 0.077 #24 of 79 Archive leaderboard report
Monocular Depth Estimation KITTI Eigen split PixelFormer Sq Rel 0.149 #24 of 79 Archive leaderboard report
Monocular Depth Estimation KITTI Eigen split PixelFormer absolute relative error 0.051 #24 of 79 Archive leaderboard report
Monocular Depth Estimation NYU-Depth V2 PixelFormer Delta < 1.25 0.929 #38 of 85 Archive leaderboard report
Monocular Depth Estimation NYU-Depth V2 PixelFormer Delta < 1.25^2 0.991 #38 of 85 Archive leaderboard report
Monocular Depth Estimation NYU-Depth V2 PixelFormer Delta < 1.25^3 0.998 #38 of 85 Archive leaderboard report
Monocular Depth Estimation NYU-Depth V2 PixelFormer RMSE 0.322 #38 of 85 Archive leaderboard report
Monocular Depth Estimation NYU-Depth V2 PixelFormer absolute relative error 0.090 #38 of 85 Archive leaderboard report
Monocular Depth Estimation NYU-Depth V2 PixelFormer log 10 0.039 #38 of 85 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Convolution

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections