{"url":"/method/cascade-r-cnn","slug":"cascade-r-cnn","name":"Cascade R-CNN","full_name":"Cascade R-CNN","full_name_withheld":false,"description_markdown":"**Cascade R-CNN** is an object detection architecture that seeks to address problems with degrading performance with increased IoU thresholds (due to overfitting during training and inference-time mismatch between IoUs for which detector is optimal and the inputs). It is a multi-stage extension of the [R-CNN](https://paperswithcode.com/method/r-cnn), where detector stages deeper into the cascade are sequentially more selective against close false positives. The cascade of R-CNN stages are trained sequentially, using the output of one stage to train the next. This is motivated by the observation that the output IoU of a regressor is almost invariably better than the input IoU. \r\n\r\nCascade R-CNN does not aim to mine hard negatives. Instead, by adjusting bounding boxes, each stage aims to find a good set of close false positives for training the next stage. When operating in this manner, a sequence of detectors adapted to increasingly higher IoUs can beat the overfitting problem, and thus be effectively trained. At inference, the same cascade procedure is applied. The progressively improved hypotheses are better matched to the increasing detector quality at each stage.","description_state":"present","introduced_year":null,"introduced_by":{"title":"Cascade R-CNN: Delving into High Quality Object Detection","paper":"/paper/cascade-r-cnn-delving-into-high-quality","first_author":"Zhaowei Cai","n_authors":2,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/cascade-r-cnn-delving-into-high-quality"},"source":{"url":"http://arxiv.org/abs/1712.00726v1","title":"Cascade R-CNN: Delving into High Quality Object Detection","url_on_a_paper_host":true},"code_snippet_url":"https://github.com/open-mmlab/mmdetection/blob/588536de9905feb7f37c2c977d146a64c74ef28e/mmdet/models/detectors/cascade_rcnn.py#L6","code_snippet_url_on_a_code_host":true,"categories":[{"area":"Computer Vision","area_id":"computer-vision","collection":"Object Detection Models","url":"/methods/category/object-detection-models","pwc_aliases":[]}],"n_papers_tagged":34,"archive_num_papers":34,"papers_newest_first":[{"paper":null,"title":"AI-Driven MRI Spine Pathology Detection: A Comprehensive Deep Learning Approach for Automated Diagnosis in Diverse Clinical Settings","date":"2025-03-26","arxiv_id":"2503.20316","n_code_links":0,"syntology":null},{"paper":null,"title":"Enhancing Tree Type Detection in Forest Fire Risk Assessment: Multi-Stage Approach and Color Encoding with Forest Fire Risk Evaluation Framework for UAV Imagery","date":"2024-07-27","arxiv_id":"2407.19184","n_code_links":0,"syntology":null},{"paper":"/paper/2408-02674","title":"On Feasibility of Intent Obfuscating Attacks","date":"2024-07-22","arxiv_id":"2408.02674","n_code_links":1,"syntology":null},{"paper":null,"title":"FAD-SAR: A Novel Fishing Activity Detection System via Synthetic Aperture Radar Images Based on Deep Learning Method","date":"2024-04-28","arxiv_id":"2404.18245","n_code_links":0,"syntology":null},{"paper":"/paper/rethinking-detection-based-table-structure","title":"Rethinking Detection Based Table Structure Recognition for Visually Rich Document Images","date":"2023-12-01","arxiv_id":"2312.00699","n_code_links":1,"syntology":null},{"paper":"/paper/semi-supervised-and-long-tailed-object","title":"Semi-Supervised and Long-Tailed Object Detection with CascadeMatch","date":"2023-05-24","arxiv_id":"2305.14813","n_code_links":1,"syntology":null},{"paper":"/paper/context-aware-chart-element-detection","title":"Context-Aware Chart Element Detection","date":"2023-05-07","arxiv_id":"2305.04151","n_code_links":1,"syntology":null},{"paper":"/paper/fqdet-fast-converging-query-based-detector","title":"FQDet: Fast-converging Query-based Detector","date":"2022-10-05","arxiv_id":"2210.02318","n_code_links":2,"syntology":{"ran":1,"of":6,"unverified":5,"pointer_only":3}},{"paper":null,"title":"ComplETR: Reducing the cost of annotations for object detection in dense scenes with vision transformers","date":"2022-09-13","arxiv_id":"2209.05654","n_code_links":0,"syntology":null},{"paper":null,"title":"Lost in Compression: the Impact of Lossy Image Compression on Variable Size Object Detection within Infrared Imagery","date":"2022-05-16","arxiv_id":"2205.08002","n_code_links":0,"syntology":null},{"paper":null,"title":"Self-Normalized Density Map (SNDM) for Counting Microbiological Objects","date":"2022-03-15","arxiv_id":"2203.09474","n_code_links":0,"syntology":null},{"paper":null,"title":"Attentional Feature Refinement and Alignment Network for Aircraft Detection in SAR Imagery","date":"2022-01-18","arxiv_id":"2201.07124","n_code_links":0,"syntology":null},{"paper":null,"title":"Automatic Detection of Injection and Press Mold Parts on 2D Drawing Using Deep Neural Network","date":"2021-10-22","arxiv_id":"2110.11593","n_code_links":0,"syntology":null},{"paper":null,"title":"Operationalizing Convolutional Neural Network Architectures for Prohibited Object Detection in X-Ray Imagery","date":"2021-10-10","arxiv_id":"2110.04906","n_code_links":0,"syntology":null},{"paper":"/paper/queryinst-parallelly-supervised-mask-query","title":"Instances as Queries","date":"2021-05-05","arxiv_id":"2105.01928","n_code_links":5,"syntology":{"ran":1,"of":3,"unverified":2,"pointer_only":0}},{"paper":"/paper/opanas-one-shot-path-aggregation-network","title":"OPANAS: One-Shot Path Aggregation Network Architecture Search for Object Detection","date":"2021-03-08","arxiv_id":"2103.04507","n_code_links":1,"syntology":null},{"paper":null,"title":"Augmenting Proposals by the Detector Itself","date":"2021-01-28","arxiv_id":"2101.11789","n_code_links":0,"syntology":null},{"paper":"/paper/synet-an-ensemble-network-for-object","title":"SyNet: An Ensemble Network for Object Detection in UAV Images","date":"2020-12-23","arxiv_id":"2012.12991","n_code_links":1,"syntology":null},{"paper":null,"title":"Hierarchical Context Embedding for Region-based Object Detection","date":"2020-08-04","arxiv_id":"2008.01338","n_code_links":0,"syntology":null},{"paper":"/paper/a-solution-to-product-detection-in-densely","title":"A Solution to Product detection in Densely Packed Scenes","date":"2020-07-23","arxiv_id":"2007.11946","n_code_links":2,"syntology":null},{"paper":"/paper/seeing-without-looking-contextual-rescoring-1","title":"Seeing without Looking: Contextual Rescoring of Object Detections for AP Maximization","date":"2020-06-01","arxiv_id":null,"n_code_links":1,"syntology":null},{"paper":null,"title":"PBRnet: Pyramidal Bounding Box Refinement to Improve Object Localization Accuracy","date":"2020-03-10","arxiv_id":"2003.04541","n_code_links":0,"syntology":null},{"paper":"/paper/side-aware-boundary-localization-for-more","title":"Side-Aware Boundary Localization for More Precise Object Detection","date":"2019-12-09","arxiv_id":"1912.04260","n_code_links":3,"syntology":null},{"paper":null,"title":"IMMVP: An Efficient Daytime and Nighttime On-Road Object Detector","date":"2019-10-15","arxiv_id":"1910.06573","n_code_links":0,"syntology":null},{"paper":"/paper/cbnet-a-novel-composite-backbone-network","title":"CBNet: A Novel Composite Backbone Network Architecture for Object Detection","date":"2019-09-09","arxiv_id":"1909.03625","n_code_links":6,"syntology":{"ran":0,"of":4,"unverified":4,"pointer_only":1}},{"paper":"/paper/instaboost-boosting-instance-segmentation-via","title":"InstaBoost: Boosting Instance Segmentation via Probability Map Guided Copy-Pasting","date":"2019-08-21","arxiv_id":"1908.07801","n_code_links":3,"syntology":null},{"paper":"/paper/190807919","title":"Deep High-Resolution Representation Learning for Visual Recognition","date":"2019-08-20","arxiv_id":"1908.07919","n_code_links":42,"syntology":{"ran":3,"of":34,"unverified":31,"pointer_only":16}},{"paper":null,"title":"Rethinking Classification and Localization for Cascade R-CNN","date":"2019-07-27","arxiv_id":"1907.11914","n_code_links":0,"syntology":null},{"paper":"/paper/cascade-r-cnn-high-quality-object-detection","title":"Cascade R-CNN: High Quality Object Detection and Instance Segmentation","date":"2019-06-24","arxiv_id":"1906.09756","n_code_links":4,"syntology":null},{"paper":"/paper/spatial-group-wise-enhance-improving-semantic","title":"Spatial Group-wise Enhance: Improving Semantic Feature Learning in Convolutional Networks","date":"2019-05-23","arxiv_id":"1905.09646","n_code_links":3,"syntology":{"ran":2,"of":4,"unverified":2,"pointer_only":4}}],"papers_shown":30,"tasks":[{"task":"/task/object-detection","name":"Object Detection","papers":23},{"task":"/task/object-detection-1","name":"object-detection","papers":18},{"task":"/task/object","name":"Object","papers":9},{"task":"/task/instance-segmentation","name":"Instance Segmentation","papers":7},{"task":"/task/semantic-segmentation","name":"Semantic Segmentation","papers":5},{"task":"/task/segmentation","name":"Segmentation","papers":3},{"task":"/task/high","name":"Vocal Bursts Intensity Prediction","papers":3},{"task":"/task/data-augmentation","name":"Data Augmentation","papers":2},{"task":"/task/classification","name":"General Classification","papers":2},{"task":"/task/image-compression","name":"Image Compression","papers":2},{"task":"/task/regression-1","name":"regression","papers":2},{"task":"/task/2d-object-detection","name":"2D Object Detection","papers":1},{"task":"/task/action-detection","name":"Action Detection","papers":1},{"task":"/task/activity-detection","name":"Activity Detection","papers":1},{"task":null,"name":"Avg","papers":1},{"task":"/task/classification-1","name":"Classification","papers":1},{"task":"/task/data-compression","name":"Data Compression","papers":1},{"task":"/task/data-visualization","name":"Data Visualization","papers":1},{"task":"/task/decoder","name":"Decoder","papers":1},{"task":"/task/dense-object-detection","name":"Dense Object Detection","papers":1}],"tasks_shown":20,"n_tasks":41,"usage_by_year":[{"year":"2017","papers":1},{"year":"2018","papers":1},{"year":"2019","papers":10},{"year":"2020","papers":5},{"year":"2021","papers":5},{"year":"2022","papers":5},{"year":"2023","papers":3},{"year":"2024","papers":3},{"year":"2025","papers":1}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/cascade-r-cnn"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}