Methods › SPEED
SPEED: Separable Pyramidal Pooling EncodEr-Decoder for Real-Time Monocular Depth Estimation on Low-Resource Settings
SPEED
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene understanding and visual odometry, which are key components in autonomous and robotic systems. Approaches based on the state of the art vision transformer architectures are extremely deep and complex not suitable for real-time inference operations on edge and autonomous systems equipped with low resources (i.e. robot indoor navigation and surveillance). This paper presents SPEED, a Separable Pyramidal pooling EncodEr-Decoder architecture designed to achieve real-time frequency performances on multiple hardware platforms. The proposed model is a fast-throughput deep architecture for MDE able to obtain depth estimations with high accuracy from low resolution images using minimum hardware resources (i.e. edge devices). Our encoder-decoder model exploits two depthwise separable pyramidal pooling layers, which allow to increase the inference frequency while reducing the overall computational complexity. The proposed method performs better than other fast-throughput architectures in terms of both accuracy and frame rates, achieving real-time performances over cloud CPU, TPU and the NVIDIA Jetson TX1 on two indoor benchmarks: the NYU Depth v2 and the DIML Kinect v2 datasets.
Papers archive 2025-07-28
30 shown of 9,576, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
Developing Visual Augmented Q&A System using Scalable Vision Embedding Retrieval & Late Interaction Re-ranker 16 Jul 2025 · 1 repository · arXiv:2507.12378
-
COLI: A Hierarchical Efficient Compressor for Large Images 15 Jul 2025 · 0 repositories · arXiv:2507.11443
-
Interpretable Bayesian Tensor Network Kernel Machines with Automatic Rank and Feature Selection 15 Jul 2025 · 1 repository · arXiv:2507.11136
-
Neurosymbolic Reasoning Shortcuts under the Independence Assumption 15 Jul 2025 · 1 repository · arXiv:2507.11357Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)
-
Streaming 4D Visual Geometry Transformer 15 Jul 2025 · 1 repository · arXiv:2507.11539Syntology ran 4 of 13 samples · 9 unverified · 13 pointer-only (licence)
-
Tomato Multi-Angle Multi-Pose Dataset for Fine-Grained Phenotyping 15 Jul 2025 · 0 repositories · arXiv:2507.11279
-
Federated Learning with Graph-Based Aggregation for Traffic Forecasting 13 Jul 2025 · 0 repositories · arXiv:2507.09805
-
Lizard: An Efficient Linearization Framework for Large Language Models 11 Jul 2025 · 0 repositories · arXiv:2507.09025
-
GNN-CNN: An Efficient Hybrid Model of Convolutional and Graph Neural Networks for Text Representation 10 Jul 2025 · 1 repository · arXiv:2507.07414
-
GSVR: 2D Gaussian-based Video Representation for 800+ FPS with Hybrid Deformation Field 8 Jul 2025 · 0 repositories · arXiv:2507.05594
-
Hyperspectral Anomaly Detection Methods: A Survey and Comparative Study 8 Jul 2025 · 0 repositories · arXiv:2507.05730
-
Robust One-step Speech Enhancement via Consistency Distillation 8 Jul 2025 · 1 repository · arXiv:2507.05688
-
Acquiring and Adapting Priors for Novel Tasks via Neural Meta-Architectures 7 Jul 2025 · 0 repositories · arXiv:2507.10446
-
MambaFusion: Height-Fidelity Dense Global Fusion for Multi-modal 3D Object Detection 6 Jul 2025 · 1 repository · arXiv:2507.04369Syntology ran 0 of 10 samples · 10 unverified
-
OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference 5 Jul 2025 · 0 repositories · arXiv:2507.03865
-
Hita: Holistic Tokenizer for Autoregressive Image Generation 3 Jul 2025 · 0 repositories · arXiv:2507.02358Syntology ran 4 of 11 samples · 7 unverified
-
CycleVAR: Repurposing Autoregressive Model for Unsupervised One-Step Image Translation 29 Jun 2025 · 1 repository · arXiv:2506.23347
-
Deterministic Object Pose Confidence Region Estimation 28 Jun 2025 · 0 repositories · arXiv:2506.22720
-
BitMark for Infinity: Watermarking Bitwise Autoregressive Image Generative Models 26 Jun 2025 · 0 repositories · arXiv:2506.21209Syntology ran 2 of 2 samples · 0 unverified
-
CaloHadronic: a diffusion model for the generation of hadronic showers 26 Jun 2025 · 1 repository · arXiv:2506.21720
-
Detection of Breast Cancer Lumpectomy Margin with SAM-incorporated Forward-Forward Contrastive Learning 26 Jun 2025 · 1 repository · arXiv:2506.21006
-
DiLoCoX: A Low-Communication Large-Scale Training Framework for Decentralized Cluster 26 Jun 2025 · 0 repositories · arXiv:2506.21263
-
ESMStereo: Enhanced ShuffleMixer Disparity Upsampling for Real-Time and Accurate Stereo Matching 26 Jun 2025 · 1 repository · arXiv:2506.21091
-
Flow-Based Single-Step Completion for Efficient and Expressive Policy Learning 26 Jun 2025 · 0 repositories · arXiv:2506.21427Syntology ran 3 of 4 samples · 1 unverified · 4 pointer-only (licence)
-
Instella-T2I: Pushing the Limits of 1D Discrete Latent Space Image Generation 26 Jun 2025 · 0 repositories · arXiv:2506.21022
-
Integrating Vehicle Acoustic Data for Enhanced Urban Traffic Management: A Study on Speed Classification in Suzhou 26 Jun 2025 · 0 repositories · arXiv:2506.21269
-
SAM4D: Segment Anything in Camera and LiDAR Streams 26 Jun 2025 · 0 repositories · arXiv:2506.21547
-
Collaborative Batch Size Optimization for Federated Learning 25 Jun 2025 · 0 repositories · arXiv:2506.20511
-
Efficient Federated Learning with Encrypted Data Sharing for Data-Heterogeneous Edge Devices 25 Jun 2025 · 0 repositories · arXiv:2506.20644
-
Exploiting Lightweight Hierarchical ViT and Dynamic Framework for Efficient Visual Tracking 25 Jun 2025 · 1 repository · arXiv:2506.20381
Tasks archive 2025-07-28
20 shown of 1,515 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
| Task | Papers |
|---|---|
| GPU | 533 |
| Object Detection | 437 |
| object-detection | 416 |
| Semantic Segmentation | 327 |
| Reinforcement Learning | 316 |
| Reinforcement Learning (RL) | 309 |
| reinforcement-learning | 303 |
| Object | 298 |
| Language Modelling | 261 |
| Decoder | 239 |
| Segmentation | 238 |
| Autonomous Driving | 225 |
| Computational Efficiency | 221 |
| Language Modeling | 219 |
| Quantization | 216 |
| CPU | 215 |
| Image Classification | 190 |
| Retrieval | 188 |
| General Classification | 184 |
| Denoising | 172 |
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
The archive places this method in no collection.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections