{"url":"/method/mask-r-cnn","slug":"mask-r-cnn","name":"Mask R-CNN","full_name":"Mask R-CNN","full_name_withheld":false,"description_markdown":"**Mask R-CNN** extends [Faster R-CNN](http://paperswithcode.com/method/faster-r-cnn) to solve instance segmentation tasks. It achieves this by adding a branch for predicting an object mask in parallel with the existing branch for bounding box recognition. In principle, Mask R-CNN is an intuitive extension of Faster [R-CNN](https://paperswithcode.com/method/r-cnn), but constructing the mask branch properly is critical for good results. \r\n\r\nMost importantly, Faster R-CNN was not designed for pixel-to-pixel alignment between network inputs and outputs. This is evident in how [RoIPool](http://paperswithcode.com/method/roi-pooling), the *de facto* core operation for attending to instances, performs coarse spatial quantization for feature extraction. To fix the misalignment, Mask R-CNN utilises a simple, quantization-free layer, called [RoIAlign](http://paperswithcode.com/method/roi-align), that faithfully preserves exact spatial locations. \r\n\r\nSecondly, Mask R-CNN *decouples* mask and class prediction: it predicts a binary mask for each class independently, without competition among classes, and relies on the network's RoI classification branch to predict the category. In contrast, an [FCN](http://paperswithcode.com/method/fcn) usually perform per-pixel multi-class categorization, which couples segmentation and classification.","description_state":"present","introduced_year":null,"introduced_by":{"title":"Mask R-CNN","paper":"/paper/mask-r-cnn","first_author":"Kaiming He","n_authors":4,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/mask-r-cnn"},"source":{"url":"http://arxiv.org/abs/1703.06870v3","title":"Mask R-CNN","url_on_a_paper_host":true},"code_snippet_url":"https://github.com/facebookresearch/detectron2/blob/601d7666faaf7eb0ba64c9f9ce5811b13861fe12/detectron2/modeling/roi_heads/mask_head.py#L154","code_snippet_url_on_a_code_host":true,"categories":[{"area":"Computer Vision","area_id":"computer-vision","collection":"Object Detection Models","url":"/methods/category/object-detection-models","pwc_aliases":[]},{"area":"Computer Vision","area_id":"computer-vision","collection":"Instance Segmentation Models","url":"/methods/category/instance-segmentation-models","pwc_aliases":[]}],"n_papers_tagged":420,"archive_num_papers":420,"papers_newest_first":[{"paper":null,"title":"A novel visual data-based diagnostic approach for estimation of regime transition in pool boiling","date":"2025-06-12","arxiv_id":"2506.10832","n_code_links":0,"syntology":null},{"paper":null,"title":"Bringing SAM to new heights: Leveraging elevation data for tree crown segmentation from drone imagery","date":"2025-06-05","arxiv_id":"2506.04970","n_code_links":0,"syntology":null},{"paper":null,"title":"Detailed Evaluation of Modern Machine Learning Approaches for Optic Plastics Sorting","date":"2025-05-22","arxiv_id":"2505.16513","n_code_links":0,"syntology":null},{"paper":null,"title":"SurgPose: Generalisable Surgical Instrument Pose Estimation using Zero-Shot Learning and Stereo Vision","date":"2025-05-16","arxiv_id":"2505.11439","n_code_links":0,"syntology":null},{"paper":null,"title":"A Robust Deep Networks based Multi-Object MultiCamera Tracking System for City Scale Traffic","date":"2025-05-01","arxiv_id":"2505.00534","n_code_links":0,"syntology":null},{"paper":null,"title":"Transcending Dimensions using Generative AI: Real-Time 3D Model Generation in Augmented Reality","date":"2025-04-27","arxiv_id":"2504.21033","n_code_links":0,"syntology":null},{"paper":null,"title":"Real-time Seafloor Segmentation and Mapping","date":"2025-04-14","arxiv_id":"2504.10750","n_code_links":0,"syntology":null},{"paper":null,"title":"RipVIS: Rip Currents Video Instance Segmentation Benchmark for Beach Monitoring and Safety","date":"2025-04-01","arxiv_id":"2504.01128","n_code_links":0,"syntology":null},{"paper":"/paper/ai-assisted-colonoscopy-polyp-detection-and","title":"AI-Assisted Colonoscopy: Polyp Detection and Segmentation using Foundation Models","date":"2025-03-31","arxiv_id":"2503.24138","n_code_links":1,"syntology":null},{"paper":null,"title":"Assessing SAM for Tree Crown Instance Segmentation from Drone Imagery","date":"2025-03-26","arxiv_id":"2503.20199","n_code_links":0,"syntology":null},{"paper":"/paper/overlock-an-overview-first-look-closely-next","title":"OverLoCK: An Overview-first-Look-Closely-next ConvNet with Context-Mixing Dynamic Kernels","date":"2025-02-27","arxiv_id":"2502.20087","n_code_links":1,"syntology":{"ran":0,"of":3,"unverified":3,"pointer_only":0}},{"paper":"/paper/sasvi-segment-any-surgical-video","title":"SASVi - Segment Any Surgical Video","date":"2025-02-12","arxiv_id":"2502.09653","n_code_links":1,"syntology":null},{"paper":null,"title":"Adaptive Object Detection for Indoor Navigation Assistance: A Performance Evaluation of Real-Time Algorithms","date":"2025-01-30","arxiv_id":"2501.18444","n_code_links":0,"syntology":null},{"paper":null,"title":"Transfer Learning for Keypoint Detection in Low-Resolution Thermal TUG Test Images","date":"2025-01-30","arxiv_id":"2501.18453","n_code_links":0,"syntology":null},{"paper":null,"title":"Effective Defect Detection Using Instance Segmentation for NDI","date":"2025-01-24","arxiv_id":"2501.14149","n_code_links":0,"syntology":null},{"paper":null,"title":"Data-driven Detection and Evaluation of Damages in Concrete Structures: Using Deep Learning and Computer Vision","date":"2025-01-21","arxiv_id":"2501.11836","n_code_links":0,"syntology":null},{"paper":"/paper/rapid-automated-mapping-of-clouds-on-titan","title":"Rapid Automated Mapping of Clouds on Titan With Instance Segmentation","date":"2025-01-08","arxiv_id":"2501.04459","n_code_links":1,"syntology":null},{"paper":null,"title":"AI-Powered Cow Detection in Complex Farm Environments","date":"2025-01-03","arxiv_id":"2501.02080","n_code_links":0,"syntology":null},{"paper":null,"title":"Leveraging Deep Learning with Multi-Head Attention for Accurate Extraction of Medicine from Handwritten Prescriptions","date":"2024-12-24","arxiv_id":"2412.18199","n_code_links":0,"syntology":null},{"paper":null,"title":"Exploring Machine Learning Engineering for Object Detection and Tracking by Unmanned Aerial Vehicle (UAV)","date":"2024-12-19","arxiv_id":"2412.15347","n_code_links":0,"syntology":null},{"paper":null,"title":"Automating Sonologists USG Commands with AI and Voice Interface","date":"2024-11-20","arxiv_id":"2411.13006","n_code_links":0,"syntology":null},{"paper":null,"title":"Raspberry PhenoSet: A Phenology-based Dataset for Automated Growth Detection and Yield Estimation","date":"2024-11-01","arxiv_id":"2411.00967","n_code_links":0,"syntology":null},{"paper":null,"title":"How Important are Data Augmentations to Close the Domain Gap for Object Detection in Orbit?","date":"2024-10-21","arxiv_id":"2410.15766","n_code_links":0,"syntology":null},{"paper":null,"title":"Segmenting objects with Bayesian fusion of active contour models and convnet priors","date":"2024-10-09","arxiv_id":"2410.07421","n_code_links":0,"syntology":null},{"paper":"/paper/synco-synthetic-hard-negatives-in-contrastive","title":"SynCo: Synthetic Hard Negatives in Contrastive Learning for Better Unsupervised Visual Representations","date":"2024-10-03","arxiv_id":"2410.02401","n_code_links":1,"syntology":null},{"paper":null,"title":"Drone Stereo Vision for Radiata Pine Branch Detection and Distance Measurement: Integrating SGBM and Segmentation Models","date":"2024-09-26","arxiv_id":"2409.17526","n_code_links":0,"syntology":null},{"paper":null,"title":"Deep-Learning Recognition of Scanning Transmission Electron Microscopy: Quantifying and Mitigating the Influence of Gaussian Noises","date":"2024-09-25","arxiv_id":"2409.16637","n_code_links":0,"syntology":null},{"paper":null,"title":"Transient Adversarial 3D Projection Attacks on Object Detection in Autonomous Driving","date":"2024-09-25","arxiv_id":"2409.17403","n_code_links":0,"syntology":null},{"paper":"/paper/low-cost-tree-crown-dieback-estimation-using","title":"Low-Cost Tree Crown Dieback Estimation Using Deep Learning-Based Segmentation","date":"2024-09-12","arxiv_id":"2409.08171","n_code_links":1,"syntology":null},{"paper":null,"title":"Knowledge Discovery in Optical Music Recognition: Enhancing Information Retrieval with Instance Segmentation","date":"2024-08-27","arxiv_id":"2408.15002","n_code_links":0,"syntology":null}],"papers_shown":30,"tasks":[{"task":"/task/semantic-segmentation","name":"Semantic Segmentation","papers":197},{"task":"/task/instance-segmentation","name":"Instance Segmentation","papers":177},{"task":"/task/object-detection","name":"Object Detection","papers":139},{"task":"/task/segmentation","name":"Segmentation","papers":135},{"task":"/task/object-detection-1","name":"object-detection","papers":121},{"task":"/task/object","name":"Object","papers":76},{"task":"/task/image-classification","name":"Image Classification","papers":27},{"task":"/task/transfer-learning","name":"Transfer Learning","papers":22},{"task":"/task/image-segmentation","name":"Image Segmentation","papers":16},{"task":"/task/pose-estimation","name":"Pose Estimation","papers":16},{"task":"/task/data-augmentation","name":"Data Augmentation","papers":15},{"task":"/task/classification","name":"General Classification","papers":14},{"task":"/task/panoptic-segmentation","name":"Panoptic Segmentation","papers":14},{"task":"/task/image-classification","name":"image-classification","papers":14},{"task":"/task/region-proposal","name":"Region Proposal","papers":12},{"task":"/task/deep-learning","name":"Deep Learning","papers":11},{"task":"/task/autonomous-driving","name":"Autonomous Driving","papers":9},{"task":"/task/clustering","name":"Clustering","papers":9},{"task":"/task/decoder","name":"Decoder","papers":9},{"task":"/task/object-recognition","name":"Object Recognition","papers":9}],"tasks_shown":20,"n_tasks":308,"usage_by_year":[{"year":"2017","papers":7},{"year":"2018","papers":27},{"year":"2019","papers":74},{"year":"2020","papers":85},{"year":"2021","papers":95},{"year":"2022","papers":38},{"year":"2023","papers":46},{"year":"2024","papers":30},{"year":"2025","papers":18}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/mask-r-cnn"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}