Methods › SPEED

SPEED: Separable Pyramidal Pooling EncodEr-Decoder for Real-Time Monocular Depth Estimation on Low-Resource Settings

SPEED

9,576 papers tagged archive 2025-07-28

archive 2025-07-28 Description, source and code snippet are the archive's method entry.

The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene understanding and visual odometry, which are key components in autonomous and robotic systems. Approaches based on the state of the art vision transformer architectures are extremely deep and complex not suitable for real-time inference operations on edge and autonomous systems equipped with low resources (i.e. robot indoor navigation and surveillance). This paper presents SPEED, a Separable Pyramidal pooling EncodEr-Decoder architecture designed to achieve real-time frequency performances on multiple hardware platforms. The proposed model is a fast-throughput deep architecture for MDE able to obtain depth estimations with high accuracy from low resolution images using minimum hardware resources (i.e. edge devices). Our encoder-decoder model exploits two depthwise separable pyramidal pooling layers, which allow to increase the inference frequency while reducing the overall computational complexity. The proposed method performs better than other fast-throughput architectures in terms of both accuracy and frame rates, achieving real-time performances over cloud CPU, TPU and the NVIDIA Jetson TX1 on two indoor benchmarks: the NYU Depth v2 and the DIML Kinect v2 datasets.

See Code · lorenzopapa5/SPEED

Papers archive 2025-07-28

30 shown of 9,576, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.

Tasks archive 2025-07-28

20 shown of 1,515 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.

TaskPapers
GPU533
Object Detection437
object-detection416
Semantic Segmentation327
Reinforcement Learning316
Reinforcement Learning (RL)309
reinforcement-learning303
Object298
Language Modelling261
Decoder239
Segmentation238
Autonomous Driving225
Computational Efficiency221
Language Modeling219
Quantization216
CPU215
Image Classification190
Retrieval188
General Classification184
Denoising172

Usage over time archive 2025-07-28

Papers per year tagged with SPEED: 2007 to 2025, peak 2,215 2,215 0 2007: 4 papers 2007 2008: 9 papers 2009: 2 papers 2009 2010: 11 papers 2011: 9 papers 2011 2012: 25 papers 2013: 67 papers 2013 2014: 140 papers 2015: 175 papers 2015 2016: 228 papers 2017: 360 papers 2017 2018: 565 papers 2019: 837 papers 2019 2020: 546 papers 2021: 326 papers 2021 2022: 1417 papers 2023: 1679 papers 2023 2024: 2215 papers 2025: 961 papers 2025
Papers per year the archive tags with this method, by the paper's archive date (9,576 dated). Bars are counts, not a trend claim.

Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).

Categories archive 2025-07-28

The archive places this method in no collection.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections