{"url":"/method/detr","slug":"detr","name":"Detr","full_name":"Detection Transformer","full_name_withheld":false,"description_markdown":"**Detr**, or **Detection Transformer**, is a set-based object detector using a [Transformer](https://paperswithcode.com/method/transformer) on top of a convolutional backbone. It uses a conventional CNN backbone to learn a 2D representation of an input image. The model flattens it and supplements it with a positional encoding before passing it into a transformer encoder. A transformer decoder then takes as input a small fixed number of learned positional embeddings, which we call object queries, and additionally attends to the encoder output. We pass each output embedding of the decoder to a shared feed forward network (FFN) that predicts either a detection (class\r\nand bounding box) or a “no object” class.","description_state":"present","introduced_year":null,"introduced_by":{"title":null,"paper":null,"first_author":null,"n_authors":0,"url_abs":null,"archive_paper_url":null},"source":{"url":"https://arxiv.org/abs/2005.12872v3","title":"End-to-End Object Detection with Transformers","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Computer Vision","area_id":"computer-vision","collection":"Object Detection Models","url":"/methods/category/object-detection-models","pwc_aliases":[]}],"n_papers_tagged":222,"archive_num_papers":null,"papers_newest_first":[{"paper":null,"title":"CF-DETR: Coarse-to-Fine Transformer for Real-Time Object Detection","date":"2025-05-29","arxiv_id":"2505.23317","n_code_links":0,"syntology":null},{"paper":null,"title":"LiDAR MOT-DETR: A LiDAR-based Two-Stage Transformer for 3D Multiple Object Tracking","date":"2025-05-19","arxiv_id":"2505.12753","n_code_links":0,"syntology":null},{"paper":null,"title":"Underwater object detection in sonar imagery with detection transformer and Zero-shot neural architecture search","date":"2025-05-10","arxiv_id":"2505.06694","n_code_links":0,"syntology":null},{"paper":null,"title":"An Efficient Aerial Image Detection with Variable Receptive Fields","date":"2025-04-21","arxiv_id":"2504.15165","n_code_links":0,"syntology":null},{"paper":"/paper/annopage-dataset-dataset-of-non-textual","title":"AnnoPage Dataset: Dataset of Non-Textual Elements in Documents with Fine-Grained Categorization","date":"2025-03-28","arxiv_id":"2503.22526","n_code_links":0,"syntology":null},{"paper":"/paper/bibliopage-a-dataset-of-scanned-title-pages","title":"BiblioPage: A Dataset of Scanned Title Pages for Bibliographic Metadata Extraction","date":"2025-03-25","arxiv_id":"2503.19658","n_code_links":1,"syntology":null},{"paper":null,"title":"LGI-DETR: Local-Global Interaction for UAV Object Detection","date":"2025-03-24","arxiv_id":"2503.18785","n_code_links":0,"syntology":null},{"paper":"/paper/fastmap-fast-queries-initialization-based","title":"FastMap: Fast Queries Initialization Based Vectorized HD Map Reconstruction Framework","date":"2025-03-07","arxiv_id":"2503.05492","n_code_links":1,"syntology":null},{"paper":"/paper/walnutdata-a-uav-remote-sensing-dataset-of","title":"WalnutData: A UAV Remote Sensing Dataset of Green Walnuts and Model Evaluation","date":"2025-02-27","arxiv_id":"2502.20092","n_code_links":1,"syntology":null},{"paper":null,"title":"Automatic Vehicle Detection using DETR: A Transformer-Based Approach for Navigating Treacherous Roads","date":"2025-02-25","arxiv_id":"2502.17843","n_code_links":0,"syntology":null},{"paper":null,"title":"RAPTOR: Refined Approach for Product Table Object Recognition","date":"2025-02-19","arxiv_id":"2502.14918","n_code_links":0,"syntology":null},{"paper":"/paper/rt-demt-a-hybrid-real-time-acupoint-detection","title":"RT-DEMT: A hybrid real-time acupoint detection model combining mamba and transformer","date":"2025-02-16","arxiv_id":"2502.11179","n_code_links":1,"syntology":null},{"paper":null,"title":"CLoCKDistill: Consistent Location-and-Context-aware Knowledge Distillation for DETRs","date":"2025-02-15","arxiv_id":"2502.10683","n_code_links":0,"syntology":null},{"paper":null,"title":"Dense Object Detection Based on De-homogenized Queries","date":"2025-02-11","arxiv_id":"2502.07194","n_code_links":0,"syntology":null},{"paper":"/paper/contourformer-real-time-contour-based-end-to","title":"ContourFormer:Real-Time Contour-Based End-to-End Instance Segmentation Transformer","date":"2025-01-29","arxiv_id":"2501.17688","n_code_links":1,"syntology":null},{"paper":null,"title":"Object Detection for Medical Image Analysis: Insights from the RT-DETR Model","date":"2025-01-27","arxiv_id":"2501.16469","n_code_links":0,"syntology":null},{"paper":null,"title":"Boosting Point-Supervised Temporal Action Localization through Integrating Query Reformation and Optimal Transport","date":"2025-01-01","arxiv_id":null,"n_code_links":0,"syntology":null},{"paper":"/paper/transformer-based-wireless-capsule-endoscopy","title":"Transformer-Based Wireless Capsule Endoscopy Bleeding Tissue Detection and Classification","date":"2024-12-26","arxiv_id":"2412.19218","n_code_links":1,"syntology":null},{"paper":null,"title":"Evaluating the Adversarial Robustness of Detection Transformers","date":"2024-12-25","arxiv_id":"2412.18718","n_code_links":0,"syntology":null},{"paper":null,"title":"Object Detection Approaches to Identifying Hand Images with High Forensic Values","date":"2024-12-21","arxiv_id":"2412.16431","n_code_links":0,"syntology":null},{"paper":"/paper/deim-detr-with-improved-matching-for-fast","title":"DEIM: DETR with Improved Matching for Fast Convergence","date":"2024-12-05","arxiv_id":"2412.04234","n_code_links":1,"syntology":{"ran":1,"of":13,"unverified":12,"pointer_only":13}},{"paper":null,"title":"Identifying Reliable Predictions in Detection Transformers","date":"2024-12-02","arxiv_id":"2412.01782","n_code_links":0,"syntology":null},{"paper":null,"title":"Excretion Detection in Pigsties Using Convolutional and Transformerbased Deep Neural Networks","date":"2024-11-29","arxiv_id":"2412.00256","n_code_links":0,"syntology":null},{"paper":null,"title":"SAT-HMR: Real-Time Multi-Person 3D Mesh Estimation via Scale-Adaptive Tokens","date":"2024-11-29","arxiv_id":"2411.19824","n_code_links":0,"syntology":null},{"paper":null,"title":"A Real-Time DETR Approach to Bangladesh Road Object Detection for Autonomous Vehicles","date":"2024-11-22","arxiv_id":"2411.15110","n_code_links":0,"syntology":null},{"paper":"/paper/exploiting-vlm-localizability-and-semantics","title":"Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection","date":"2024-11-17","arxiv_id":"2411.10922","n_code_links":1,"syntology":null},{"paper":"/paper/retr-multi-view-radar-detection-transformer","title":"RETR: Multi-View Radar Detection Transformer for Indoor Perception","date":"2024-11-15","arxiv_id":"2411.10293","n_code_links":1,"syntology":{"ran":2,"of":13,"unverified":11,"pointer_only":13}},{"paper":"/paper/d-fine-redefine-regression-task-in-detrs-as","title":"D-FINE: Redefine Regression Task in DETRs as Fine-grained Distribution Refinement","date":"2024-10-17","arxiv_id":"2410.13842","n_code_links":5,"syntology":{"ran":10,"of":14,"unverified":4,"pointer_only":0}},{"paper":"/paper/unig-modelling-unitary-3d-gaussians-for-view","title":"UniGS: Modeling Unitary 3D Gaussians for Novel View Synthesis from Sparse-view Images","date":"2024-10-17","arxiv_id":"2410.13195","n_code_links":2,"syntology":null},{"paper":null,"title":"MambaBEV: An efficient 3D detection model with Mamba2","date":"2024-10-16","arxiv_id":"2410.12673","n_code_links":0,"syntology":null}],"papers_shown":30,"tasks":[{"task":"/task/object-detection","name":"Object Detection","papers":143},{"task":"/task/object-detection-1","name":"object-detection","papers":125},{"task":"/task/object","name":"Object","papers":72},{"task":"/task/decoder","name":"Decoder","papers":53},{"task":"/task/semantic-segmentation","name":"Semantic Segmentation","papers":16},{"task":"/task/instance-segmentation","name":"Instance Segmentation","papers":14},{"task":"/task/real-time-object-detection","name":"Real-Time Object Detection","papers":10},{"task":"/task/segmentation","name":"Segmentation","papers":10},{"task":"/task/autonomous-driving","name":"Autonomous Driving","papers":9},{"task":"/task/knowledge-distillation","name":"Knowledge Distillation","papers":9},{"task":"/task/2d-object-detection","name":"2D Object Detection","papers":8},{"task":"/task/contrastive-learning","name":"Contrastive Learning","papers":8},{"task":"/task/data-augmentation","name":"Data Augmentation","papers":7},{"task":null,"name":"GPU","papers":7},{"task":"/task/image-classification","name":"Image Classification","papers":7},{"task":"/task/few-shot-object-detection","name":"Few-Shot Object Detection","papers":6},{"task":"/task/action-detection","name":"Action Detection","papers":5},{"task":"/task/language-modelling","name":"Language Modelling","papers":5},{"task":"/task/moment-retrieval","name":"Moment Retrieval","papers":5},{"task":"/task/multi-object-tracking","name":"Multi-Object Tracking","papers":5}],"tasks_shown":20,"n_tasks":162,"usage_by_year":[{"year":"2020","papers":8},{"year":"2021","papers":37},{"year":"2022","papers":39},{"year":"2023","papers":57},{"year":"2024","papers":64},{"year":"2025","papers":17}],"row_source":"embedded","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/detr"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}