{"url":"/method/sparse-r-cnn","slug":"sparse-r-cnn","name":"Sparse R-CNN","full_name":"Sparse R-CNN","full_name_withheld":false,"description_markdown":"**Sparse R-CNN** is a purely sparse method for object detection in images, without object positional candidates enumerating\r\non all(dense) image grids nor object queries interacting with global(dense) image feature.\r\n\r\nAs shown in the Figure, object candidates are given with a fixed small set of learnable bounding boxes represented by 4-d coordinate. For the example of the COCO dataset, 100 boxes and 400 parameters are needed in total, rather than the predicted ones from hundreds of thousands of candidates in a Region Proposal Network ([RPN](https://paperswithcode.com/method/rpn)). These sparse candidates are used as proposal boxes to extract the feature of Region of Interest (RoI) by [RoIPool](https://paperswithcode.com/method/roi-pooling) or [RoIAlign](https://paperswithcode.com/method/roi-align).","description_state":"present","introduced_year":null,"introduced_by":{"title":null,"paper":null,"first_author":null,"n_authors":0,"url_abs":null,"archive_paper_url":null},"source":{"url":"https://arxiv.org/abs/2011.12450v2","title":"Sparse R-CNN: End-to-End Object Detection with Learnable Proposals","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Computer Vision","area_id":"computer-vision","collection":"Object Detection Models","url":"/methods/category/object-detection-models","pwc_aliases":[]}],"n_papers_tagged":16,"archive_num_papers":null,"papers_newest_first":[{"paper":"/paper/sparse-r-cnn-obb-ship-target-detection-in-sar","title":"Sparse R-CNN OBB: Ship Target Detection in SAR Images Based on Oriented Sparse Proposals","date":"2024-09-12","arxiv_id":"2409.07973","n_code_links":2,"syntology":null},{"paper":null,"title":"Enhancing Tree Type Detection in Forest Fire Risk Assessment: Multi-Stage Approach and Color Encoding with Forest Fire Risk Evaluation Framework for UAV Imagery","date":"2024-07-27","arxiv_id":"2407.19184","n_code_links":0,"syntology":null},{"paper":"/paper/polyr-cnn-r-cnn-for-end-to-end-polygonal","title":"PolyR-CNN: R-CNN for end-to-end polygonal building outline extraction","date":"2024-07-20","arxiv_id":"2407.14912","n_code_links":1,"syntology":null},{"paper":null,"title":"Stability Plasticity Decoupled Fine-tuning For Few-shot end-to-end Object Detection","date":"2024-01-20","arxiv_id":"2401.11140","n_code_links":0,"syntology":null},{"paper":null,"title":"M&M: Tackling False Positives in Mammography with a Multi-view and Multi-instance Learning Sparse Detector","date":"2023-08-11","arxiv_id":"2308.06420","n_code_links":0,"syntology":null},{"paper":"/paper/recursivedet-end-to-end-region-based","title":"RecursiveDet: End-to-End Region-based Recursive Object Detection","date":"2023-07-25","arxiv_id":"2307.13619","n_code_links":1,"syntology":{"ran":5,"of":7,"unverified":2,"pointer_only":0}},{"paper":"/paper/semi-supervised-and-long-tailed-object","title":"Semi-Supervised and Long-Tailed Object Detection with CascadeMatch","date":"2023-05-24","arxiv_id":"2305.14813","n_code_links":1,"syntology":null},{"paper":"/paper/correlation-loss-enforcing-correlation","title":"Correlation Loss: Enforcing Correlation between Classification and Localization","date":"2023-01-03","arxiv_id":"2301.01019","n_code_links":1,"syntology":null},{"paper":"/paper/spatio-temporal-learnable-proposals-for-end","title":"Spatio-Temporal Learnable Proposals for End-to-End Video Object Detection","date":"2022-10-05","arxiv_id":"2210.02368","n_code_links":0,"syntology":null},{"paper":"/paper/iou-enhanced-attention-for-end-to-end-task","title":"IoU-Enhanced Attention for End-to-End Task Specific Object Detection","date":"2022-09-21","arxiv_id":"2209.10391","n_code_links":2,"syntology":null},{"paper":"/paper/srcn3d-sparse-r-cnn-3d-surround-view-camera","title":"SRCN3D: Sparse R-CNN 3D for Compact Convolutional Multi-View 3D Object Detection and Tracking","date":"2022-06-29","arxiv_id":"2206.14451","n_code_links":2,"syntology":null},{"paper":"/paper/featurized-query-r-cnn","title":"Featurized Query R-CNN","date":"2022-06-13","arxiv_id":"2206.06258","n_code_links":1,"syntology":null},{"paper":null,"title":"Dynamic Sparse R-CNN","date":"2022-05-04","arxiv_id":"2205.02101","n_code_links":0,"syntology":null},{"paper":"/paper/structured-sparse-r-cnn-for-direct-scene","title":"Structured Sparse R-CNN for Direct Scene Graph Generation","date":"2021-06-21","arxiv_id":"2106.10815","n_code_links":4,"syntology":{"ran":1,"of":1,"unverified":0,"pointer_only":1}},{"paper":"/paper/queryinst-parallelly-supervised-mask-query","title":"Instances as Queries","date":"2021-05-05","arxiv_id":"2105.01928","n_code_links":5,"syntology":{"ran":1,"of":3,"unverified":2,"pointer_only":0}},{"paper":"/paper/sparse-r-cnn-end-to-end-object-detection-with","title":"Sparse R-CNN: End-to-End Object Detection with Learnable Proposals","date":"2020-11-25","arxiv_id":"2011.12450","n_code_links":6,"syntology":null}],"papers_shown":16,"tasks":[{"task":"/task/object-detection","name":"Object Detection","papers":13},{"task":"/task/object-detection-1","name":"object-detection","papers":13},{"task":"/task/object","name":"Object","papers":10},{"task":"/task/2d-object-detection","name":"2D Object Detection","papers":1},{"task":"/task/3d-multi-object-tracking","name":"3D Multi-Object Tracking","papers":1},{"task":"/task/3d-object-detection","name":"3D Object Detection","papers":1},{"task":"/task/autonomous-driving","name":"Autonomous Driving","papers":1},{"task":"/task/classification-1","name":"Classification","papers":1},{"task":"/task/decoder","name":"Decoder","papers":1},{"task":"/task/extracting-buildings-in-remote-sensing-images","name":"Extracting Buildings In Remote Sensing Images","papers":1},{"task":"/task/few-shot-object-detection","name":"Few-Shot Object Detection","papers":1},{"task":"/task/fire-detection","name":"Fire Detection","papers":1},{"task":"/task/graph-generation","name":"Graph Generation","papers":1},{"task":"/task/inductive-bias","name":"Inductive Bias","papers":1},{"task":"/task/instance-segmentation","name":"Instance Segmentation","papers":1},{"task":"/task/knowledge-distillation","name":"Knowledge Distillation","papers":1},{"task":"/task/long-tailed-object-detection","name":"Long-tailed Object Detection","papers":1},{"task":"/task/management","name":"Management","papers":1},{"task":"/task/multi-object-tracking","name":"Multi-Object Tracking","papers":1},{"task":"/task/object-recognition","name":"Object Recognition","papers":1}],"tasks_shown":20,"n_tasks":30,"usage_by_year":[{"year":"2020","papers":1},{"year":"2021","papers":2},{"year":"2022","papers":5},{"year":"2023","papers":4},{"year":"2024","papers":4}],"row_source":"embedded","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/sparse-r-cnn"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}