{"url":"/method/ssd","slug":"ssd","name":"SSD","full_name":"SSD","full_name_withheld":false,"description_markdown":"**SSD** is a single-stage object detection method that discretizes the output space of bounding boxes into a set of default boxes over different aspect ratios and scales per feature map location. At prediction time, the network generates scores for the presence of each object category in each default box and produces adjustments to the box to better match the object shape. Additionally, the network combines predictions from multiple feature maps with different resolutions to naturally handle objects of various sizes. \r\n\r\nThe fundamental improvement in speed comes from eliminating bounding box proposals and the subsequent pixel or feature resampling stage. Improvements over competing single-stage methods include using a small convolutional filter to predict object categories and offsets in bounding box locations, using separate predictors (filters) for different aspect ratio detections, and applying these filters to multiple feature maps from the later stages of a network in order to perform detection at multiple scales.","description_state":"present","introduced_year":null,"introduced_by":{"title":"SSD: Single Shot MultiBox Detector","paper":"/paper/ssd-single-shot-multibox-detector","first_author":"Wei Liu","n_authors":7,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/ssd-single-shot-multibox-detector"},"source":{"url":"http://arxiv.org/abs/1512.02325v5","title":"SSD: Single Shot MultiBox Detector","url_on_a_paper_host":true},"code_snippet_url":"https://github.com/amdegroot/ssd.pytorch/blob/5b0b77faa955c1917b0c710d770739ba8fbff9b7/ssd.py#L10","code_snippet_url_on_a_code_host":true,"categories":[{"area":"Computer Vision","area_id":"computer-vision","collection":"One-Stage Object Detection Models","url":"/methods/category/one-stage-object-detection-models","pwc_aliases":[]},{"area":"Computer Vision","area_id":"computer-vision","collection":"Object Detection Models","url":"/methods/category/object-detection-models","pwc_aliases":[]}],"n_papers_tagged":278,"archive_num_papers":278,"papers_newest_first":[{"paper":null,"title":"ECORE: Energy-Conscious Optimized Routing for Deep Learning Models at the Edge","date":"2025-07-08","arxiv_id":"2507.06011","n_code_links":0,"syntology":null},{"paper":null,"title":"Optimization of bi-directional gated loop cell based on multi-head attention mechanism for SSD health state classification model","date":"2025-06-13","arxiv_id":"2506.14830","n_code_links":0,"syntology":null},{"paper":null,"title":"Cost-Efficient LLM Training with Lifetime-Aware Tensor Offloading via GPUDirect Storage","date":"2025-06-06","arxiv_id":"2506.06472","n_code_links":0,"syntology":null},{"paper":null,"title":"Accelerating Autoregressive Speech Synthesis Inference With Speech Speculative Decoding","date":"2025-05-21","arxiv_id":"2505.15380","n_code_links":0,"syntology":null},{"paper":null,"title":"Defect Detection in Photolithographic Patterns Using Deep Learning Models Trained on Synthetic Data","date":"2025-05-15","arxiv_id":"2505.10192","n_code_links":0,"syntology":null},{"paper":null,"title":"StableMotion: Repurposing Diffusion-Based Image Priors for Motion Estimation","date":"2025-05-10","arxiv_id":"2505.06668","n_code_links":0,"syntology":null},{"paper":null,"title":"PaniCar: Securing the Perception of Advanced Driving Assistance Systems Against Emergency Vehicle Lighting","date":"2025-05-08","arxiv_id":"2505.05183","n_code_links":0,"syntology":null},{"paper":null,"title":"Learning to Borrow Features for Improved Detection of Small Objects in Single-Shot Detectors","date":"2025-04-30","arxiv_id":"2505.00044","n_code_links":0,"syntology":null},{"paper":null,"title":"ForcePose: A Deep Learning Approach for Force Calculation Based on Action Recognition Using MediaPipe Pose Estimation Combined with Object Detection","date":"2025-03-28","arxiv_id":"2503.22363","n_code_links":0,"syntology":null},{"paper":null,"title":"Adaptive Object Detection for Indoor Navigation Assistance: A Performance Evaluation of Real-Time Algorithms","date":"2025-01-30","arxiv_id":"2501.18444","n_code_links":0,"syntology":null},{"paper":null,"title":"Object Detection for Medical Image Analysis: Insights from the RT-DETR Model","date":"2025-01-27","arxiv_id":"2501.16469","n_code_links":0,"syntology":null},{"paper":null,"title":"Variational U-Net with Local Alignment for Joint Tumor Extraction and Registration (VALOR-Net) of Breast MRI Data Acquired at Two Different Field Strengths","date":"2025-01-23","arxiv_id":"2501.13690","n_code_links":0,"syntology":null},{"paper":null,"title":"Diffusion Model is Effectively Its Own Teacher","date":"2025-01-01","arxiv_id":null,"n_code_links":0,"syntology":null},{"paper":null,"title":"Optimizing SSD Caches for Cloud Block Storage Systems Using Machine Learning Approaches","date":"2024-12-29","arxiv_id":"2501.14770","n_code_links":0,"syntology":null},{"paper":null,"title":"Deep Learning and Hybrid Approaches for Dynamic Scene Analysis, Object Detection and Motion Tracking","date":"2024-12-05","arxiv_id":"2412.05331","n_code_links":0,"syntology":null},{"paper":null,"title":"InfiniDreamer: Arbitrarily Long Human Motion Generation via Segment Score Distillation","date":"2024-11-27","arxiv_id":"2411.18303","n_code_links":0,"syntology":null},{"paper":"/paper/efficientvim-efficient-vision-mamba-with","title":"EfficientViM: Efficient Vision Mamba with Hidden State Mixer based State Space Duality","date":"2024-11-22","arxiv_id":"2411.15241","n_code_links":2,"syntology":null},{"paper":null,"title":"Harnessing Your DRAM and SSD for Sustainable and Accessible LLM Inference with Mixed-Precision and Multi-level Caching","date":"2024-10-17","arxiv_id":"2410.14740","n_code_links":0,"syntology":null},{"paper":"/paper/exploring-the-benefit-of-activation-sparsity","title":"Exploring the Benefit of Activation Sparsity in Pre-training","date":"2024-10-04","arxiv_id":"2410.03440","n_code_links":1,"syntology":{"ran":4,"of":4,"unverified":0,"pointer_only":4}},{"paper":null,"title":"Benchmarking Deep Learning Models for Object Detection on Edge Computing Devices","date":"2024-09-25","arxiv_id":"2409.16808","n_code_links":0,"syntology":null},{"paper":"/paper/tim4rec-an-efficient-sequential","title":"TiM4Rec: An Efficient Sequential Recommendation Model Based on Time-Aware Structured State Space Duality Model","date":"2024-09-24","arxiv_id":"2409.16182","n_code_links":1,"syntology":null},{"paper":null,"title":"Vision Language Model for Interpretable and Fine-grained Detection of Safety Compliance in Diverse Workplaces","date":"2024-08-13","arxiv_id":"2408.07146","n_code_links":0,"syntology":null},{"paper":"/paper/vssd-vision-mamba-with-non-casual-state-space","title":"VSSD: Vision Mamba with Non-Causal State Space Duality","date":"2024-07-26","arxiv_id":"2407.18559","n_code_links":2,"syntology":{"ran":4,"of":11,"unverified":7,"pointer_only":1}},{"paper":"/paper/2408-02674","title":"On Feasibility of Intent Obfuscating Attacks","date":"2024-07-22","arxiv_id":"2408.02674","n_code_links":1,"syntology":null},{"paper":"/paper/historical-ink-semantic-shift-detection-for","title":"Historical Ink: Semantic Shift Detection for 19th Century Spanish","date":"2024-07-08","arxiv_id":"2407.12852","n_code_links":1,"syntology":null},{"paper":null,"title":"Scarecrow monitoring system:employing mobilenet ssd for enhanced animal supervision","date":"2024-07-01","arxiv_id":"2407.01435","n_code_links":0,"syntology":null},{"paper":"/paper/mooncake-a-kvcache-centric-disaggregated","title":"Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving","date":"2024-06-24","arxiv_id":"2407.00079","n_code_links":1,"syntology":{"ran":1,"of":1,"unverified":0,"pointer_only":0}},{"paper":null,"title":"GraphEval2000: Benchmarking and Improving Large Language Models on Graph Datasets","date":"2024-06-23","arxiv_id":"2406.16176","n_code_links":0,"syntology":null},{"paper":"/paper/reducing-memory-contention-and-i-o-congestion","title":"Reducing Memory Contention and I/O Congestion for Disk-based GNN Training","date":"2024-06-20","arxiv_id":"2406.13984","n_code_links":1,"syntology":null},{"paper":null,"title":"Endor: Hardware-Friendly Sparse Format for Offloaded LLM Inference","date":"2024-06-17","arxiv_id":"2406.11674","n_code_links":0,"syntology":null}],"papers_shown":30,"tasks":[{"task":"/task/object-detection","name":"Object Detection","papers":126},{"task":"/task/object-detection-1","name":"object-detection","papers":120},{"task":"/task/object","name":"Object","papers":66},{"task":null,"name":"GPU","papers":16},{"task":"/task/real-time-object-detection","name":"Real-Time Object Detection","papers":16},{"task":"/task/image-classification","name":"Image Classification","papers":11},{"task":"/task/deep-learning","name":"Deep Learning","papers":10},{"task":"/task/pedestrian-detection","name":"Pedestrian Detection","papers":10},{"task":"/task/transfer-learning","name":"Transfer Learning","papers":10},{"task":"/task/classification","name":"General Classification","papers":8},{"task":"/task/image-classification","name":"image-classification","papers":8},{"task":"/task/autonomous-driving","name":"Autonomous Driving","papers":7},{"task":null,"name":"CPU","papers":7},{"task":"/task/semantic-segmentation","name":"Semantic Segmentation","papers":7},{"task":"/task/classification-1","name":"Classification","papers":6},{"task":"/task/regression-1","name":"regression","papers":6},{"task":"/task/2d-object-detection","name":"2D Object Detection","papers":5},{"task":"/task/benchmarking","name":"Benchmarking","papers":5},{"task":"/task/clustering","name":"Clustering","papers":5},{"task":"/task/quantization","name":"Quantization","papers":5}],"tasks_shown":20,"n_tasks":234,"usage_by_year":[{"year":"2015","papers":1},{"year":"2016","papers":6},{"year":"2017","papers":21},{"year":"2018","papers":42},{"year":"2019","papers":33},{"year":"2020","papers":36},{"year":"2021","papers":45},{"year":"2022","papers":22},{"year":"2023","papers":24},{"year":"2024","papers":35},{"year":"2025","papers":13}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/ssd"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}