{"url":"/method/speed","slug":"speed","name":"SPEED","full_name":"SPEED: Separable Pyramidal Pooling EncodEr-Decoder for Real-Time Monocular Depth Estimation on Low-Resource Settings","full_name_withheld":false,"description_markdown":"The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene understanding and visual odometry, which are key components in autonomous and robotic systems. \r\nApproaches based on the state of the art vision transformer architectures are extremely deep and complex not suitable for real-time inference operations on edge and autonomous systems equipped with low resources (i.e. robot indoor navigation and surveillance). This paper presents SPEED, a Separable Pyramidal pooling EncodEr-Decoder architecture designed to achieve real-time frequency performances on multiple hardware platforms. The proposed model is a fast-throughput deep architecture for MDE able to obtain depth estimations with high accuracy from low resolution images using minimum hardware resources (i.e. edge devices). Our encoder-decoder model exploits two depthwise separable pyramidal pooling layers, which allow to increase the inference frequency while reducing the overall computational complexity. The proposed method performs better than other fast-throughput architectures in terms of both accuracy and frame rates, achieving real-time performances over cloud CPU, TPU and the NVIDIA Jetson TX1 on two indoor benchmarks: the NYU Depth v2 and the DIML Kinect v2 datasets.","description_state":"present","introduced_year":null,"introduced_by":{"title":null,"paper":null,"first_author":null,"n_authors":0,"url_abs":null,"archive_paper_url":null},"source":{"url":null,"title":null,"url_on_a_paper_host":false},"code_snippet_url":"https://github.com/lorenzopapa5/SPEED","code_snippet_url_on_a_code_host":true,"categories":[],"n_papers_tagged":9576,"archive_num_papers":9576,"papers_newest_first":[{"paper":"/paper/developing-visual-augmented-q-a-system-using","title":"Developing Visual Augmented Q&A System using Scalable Vision Embedding Retrieval & Late Interaction Re-ranker","date":"2025-07-16","arxiv_id":"2507.12378","n_code_links":1,"syntology":null},{"paper":null,"title":"COLI: A Hierarchical Efficient Compressor for Large Images","date":"2025-07-15","arxiv_id":"2507.11443","n_code_links":0,"syntology":null},{"paper":"/paper/interpretable-bayesian-tensor-network-kernel","title":"Interpretable Bayesian Tensor Network Kernel Machines with Automatic Rank and Feature Selection","date":"2025-07-15","arxiv_id":"2507.11136","n_code_links":1,"syntology":null},{"paper":"/paper/neurosymbolic-reasoning-shortcuts-under-the","title":"Neurosymbolic Reasoning Shortcuts under the Independence Assumption","date":"2025-07-15","arxiv_id":"2507.11357","n_code_links":1,"syntology":{"ran":1,"of":1,"unverified":0,"pointer_only":1}},{"paper":"/paper/streaming-4d-visual-geometry-transformer","title":"Streaming 4D Visual Geometry Transformer","date":"2025-07-15","arxiv_id":"2507.11539","n_code_links":1,"syntology":{"ran":4,"of":13,"unverified":9,"pointer_only":13}},{"paper":null,"title":"Tomato Multi-Angle Multi-Pose Dataset for Fine-Grained Phenotyping","date":"2025-07-15","arxiv_id":"2507.11279","n_code_links":0,"syntology":null},{"paper":null,"title":"Federated Learning with Graph-Based Aggregation for Traffic Forecasting","date":"2025-07-13","arxiv_id":"2507.09805","n_code_links":0,"syntology":null},{"paper":null,"title":"Lizard: An Efficient Linearization Framework for Large Language Models","date":"2025-07-11","arxiv_id":"2507.09025","n_code_links":0,"syntology":null},{"paper":"/paper/gnn-cnn-an-efficient-hybrid-model-of-1","title":"GNN-CNN: An Efficient Hybrid Model of Convolutional and Graph Neural Networks for Text Representation","date":"2025-07-10","arxiv_id":"2507.07414","n_code_links":1,"syntology":null},{"paper":null,"title":"GSVR: 2D Gaussian-based Video Representation for 800+ FPS with Hybrid Deformation Field","date":"2025-07-08","arxiv_id":"2507.05594","n_code_links":0,"syntology":null},{"paper":null,"title":"Hyperspectral Anomaly Detection Methods: A Survey and Comparative Study","date":"2025-07-08","arxiv_id":"2507.05730","n_code_links":0,"syntology":null},{"paper":"/paper/robust-one-step-speech-enhancement-via-1","title":"Robust One-step Speech Enhancement via Consistency Distillation","date":"2025-07-08","arxiv_id":"2507.05688","n_code_links":1,"syntology":null},{"paper":null,"title":"Acquiring and Adapting Priors for Novel Tasks via Neural Meta-Architectures","date":"2025-07-07","arxiv_id":"2507.10446","n_code_links":0,"syntology":null},{"paper":"/paper/mambafusion-height-fidelity-dense-global","title":"MambaFusion: Height-Fidelity Dense Global Fusion for Multi-modal 3D Object Detection","date":"2025-07-06","arxiv_id":"2507.04369","n_code_links":1,"syntology":{"ran":0,"of":10,"unverified":10,"pointer_only":0}},{"paper":null,"title":"OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference","date":"2025-07-05","arxiv_id":"2507.03865","n_code_links":0,"syntology":null},{"paper":"/paper/hita-holistic-tokenizer-for-autoregressive","title":"Hita: Holistic Tokenizer for Autoregressive Image Generation","date":"2025-07-03","arxiv_id":"2507.02358","n_code_links":0,"syntology":{"ran":4,"of":11,"unverified":7,"pointer_only":0}},{"paper":"/paper/cyclevar-repurposing-autoregressive-model-for","title":"CycleVAR: Repurposing Autoregressive Model for Unsupervised One-Step Image Translation","date":"2025-06-29","arxiv_id":"2506.23347","n_code_links":1,"syntology":null},{"paper":null,"title":"Deterministic Object Pose Confidence Region Estimation","date":"2025-06-28","arxiv_id":"2506.22720","n_code_links":0,"syntology":null},{"paper":"/paper/bitmark-for-infinity-watermarking-bitwise","title":"BitMark for Infinity: Watermarking Bitwise Autoregressive Image Generative Models","date":"2025-06-26","arxiv_id":"2506.21209","n_code_links":0,"syntology":{"ran":2,"of":2,"unverified":0,"pointer_only":0}},{"paper":"/paper/calohadronic-a-diffusion-model-for-the","title":"CaloHadronic: a diffusion model for the generation of hadronic showers","date":"2025-06-26","arxiv_id":"2506.21720","n_code_links":1,"syntology":null},{"paper":"/paper/detection-of-breast-cancer-lumpectomy-margin","title":"Detection of Breast Cancer Lumpectomy Margin with SAM-incorporated Forward-Forward Contrastive Learning","date":"2025-06-26","arxiv_id":"2506.21006","n_code_links":1,"syntology":null},{"paper":null,"title":"DiLoCoX: A Low-Communication Large-Scale Training Framework for Decentralized Cluster","date":"2025-06-26","arxiv_id":"2506.21263","n_code_links":0,"syntology":null},{"paper":"/paper/esmstereo-enhanced-shufflemixer-disparity","title":"ESMStereo: Enhanced ShuffleMixer Disparity Upsampling for Real-Time and Accurate Stereo Matching","date":"2025-06-26","arxiv_id":"2506.21091","n_code_links":1,"syntology":null},{"paper":"/paper/flow-based-single-step-completion-for","title":"Flow-Based Single-Step Completion for Efficient and Expressive Policy Learning","date":"2025-06-26","arxiv_id":"2506.21427","n_code_links":0,"syntology":{"ran":3,"of":4,"unverified":1,"pointer_only":4}},{"paper":null,"title":"Instella-T2I: Pushing the Limits of 1D Discrete Latent Space Image Generation","date":"2025-06-26","arxiv_id":"2506.21022","n_code_links":0,"syntology":null},{"paper":null,"title":"Integrating Vehicle Acoustic Data for Enhanced Urban Traffic Management: A Study on Speed Classification in Suzhou","date":"2025-06-26","arxiv_id":"2506.21269","n_code_links":0,"syntology":null},{"paper":null,"title":"SAM4D: Segment Anything in Camera and LiDAR Streams","date":"2025-06-26","arxiv_id":"2506.21547","n_code_links":0,"syntology":null},{"paper":null,"title":"Collaborative Batch Size Optimization for Federated Learning","date":"2025-06-25","arxiv_id":"2506.20511","n_code_links":0,"syntology":null},{"paper":null,"title":"Efficient Federated Learning with Encrypted Data Sharing for Data-Heterogeneous Edge Devices","date":"2025-06-25","arxiv_id":"2506.20644","n_code_links":0,"syntology":null},{"paper":"/paper/exploiting-lightweight-hierarchical-vit-and","title":"Exploiting Lightweight Hierarchical ViT and Dynamic Framework for Efficient Visual Tracking","date":"2025-06-25","arxiv_id":"2506.20381","n_code_links":1,"syntology":null}],"papers_shown":30,"tasks":[{"task":null,"name":"GPU","papers":533},{"task":"/task/object-detection","name":"Object Detection","papers":437},{"task":"/task/object-detection-1","name":"object-detection","papers":416},{"task":"/task/semantic-segmentation","name":"Semantic Segmentation","papers":327},{"task":"/task/reinforcement-learning","name":"Reinforcement Learning","papers":316},{"task":"/task/reinforcement-learning-1","name":"Reinforcement Learning (RL)","papers":309},{"task":"/task/reinforcement-learning-2","name":"reinforcement-learning","papers":303},{"task":"/task/object","name":"Object","papers":298},{"task":"/task/language-modelling","name":"Language Modelling","papers":261},{"task":"/task/decoder","name":"Decoder","papers":239},{"task":"/task/segmentation","name":"Segmentation","papers":238},{"task":"/task/autonomous-driving","name":"Autonomous Driving","papers":225},{"task":"/task/computational-efficiency","name":"Computational Efficiency","papers":221},{"task":"/task/language-modeling","name":"Language Modeling","papers":219},{"task":"/task/quantization","name":"Quantization","papers":216},{"task":null,"name":"CPU","papers":215},{"task":"/task/image-classification","name":"Image Classification","papers":190},{"task":"/task/retrieval","name":"Retrieval","papers":188},{"task":"/task/classification","name":"General Classification","papers":184},{"task":"/task/denoising","name":"Denoising","papers":172}],"tasks_shown":20,"n_tasks":1515,"usage_by_year":[{"year":"2007","papers":4},{"year":"2008","papers":9},{"year":"2009","papers":2},{"year":"2010","papers":11},{"year":"2011","papers":9},{"year":"2012","papers":25},{"year":"2013","papers":67},{"year":"2014","papers":140},{"year":"2015","papers":175},{"year":"2016","papers":228},{"year":"2017","papers":360},{"year":"2018","papers":565},{"year":"2019","papers":837},{"year":"2020","papers":546},{"year":"2021","papers":326},{"year":"2022","papers":1417},{"year":"2023","papers":1679},{"year":"2024","papers":2215},{"year":"2025","papers":961}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/speed"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}