Papers › ViDT: An Efficient and Effective Fully Transformer-based Object Detector

ViDT: An Efficient and Effective Fully Transformer-based Object Detector

8 Oct 2021ICLR 2022 4arXiv:2110.03921archive 2025-07-28

Hwanjun Song, Deqing Sun, Sanghyuk Chun, Varun Jampani, Dongyoon Han, Byeongho Heo, Wonjae Kim, Ming-Hsuan Yang

Transformers are transforming the landscape of computer vision, especially for recognition tasks. Detection transformers are the first fully end-to-end learning systems for object detection, while vision transformers are the first fully transformer-based architecture for image classification. In this paper, we integrate Vision and Detection Transformers (ViDT) to build an effective and efficient object detector. ViDT introduces a reconfigured attention module to extend the recent Swin Transformer to be a standalone object detector, followed by a computationally efficient transformer decoder that exploits multi-scale features and auxiliary techniques essential to boost the detection performance without much increase in computational load. Extensive evaluation results on the Microsoft COCO benchmark dataset demonstrate that ViDT obtains the best AP and latency trade-off among existing fully transformer-based object detectors, and achieves 49.2AP owing to its high scalability for large models. We will release the code and trained models at https://github.com/naver-ai/vidt

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

naver-ai/vidt officialmentioned in papermentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

DecoderImage ClassificationObjectObject Detectionimage-classificationobject-detection

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Object Detection COCO 2017 val ViDT Swin-base AP 49.2 #19 of 33 Archive leaderboard report
Object Detection COCO 2017 val ViDT Swin-base AP50 69.4 #19 of 33 Archive leaderboard report
Object Detection COCO 2017 val ViDT Swin-base AP75 53.1 #19 of 33 Archive leaderboard report
Object Detection COCO 2017 val ViDT Swin-base APL 66.9 #19 of 33 Archive leaderboard report
Object Detection COCO 2017 val ViDT Swin-base APM 52.6 #19 of 33 Archive leaderboard report
Object Detection COCO 2017 val ViDT Swin-base APS 30.6 #19 of 33 Archive leaderboard report
Object Detection COCO 2017 val ViDT Swin-base Param. 0.1B #19 of 33 Archive leaderboard report
Object Detection COCO 2017 val ViDT Swin-small AP 47.5 #24 of 33 Archive leaderboard report
Object Detection COCO 2017 val ViDT Swin-small AP50 67.7 #24 of 33 Archive leaderboard report
Object Detection COCO 2017 val ViDT Swin-small AP75 51.4 #24 of 33 Archive leaderboard report
Object Detection COCO 2017 val ViDT Swin-small APL 64.8 #24 of 33 Archive leaderboard report
Object Detection COCO 2017 val ViDT Swin-small APM 50.7 #24 of 33 Archive leaderboard report
Object Detection COCO 2017 val ViDT Swin-small APS 29.2 #24 of 33 Archive leaderboard report
Object Detection COCO 2017 val ViDT Swin-small Param. 61M #24 of 33 Archive leaderboard report
Object Detection COCO 2017 val ViDT Swin-tiny AP 44.8 #27 of 33 Archive leaderboard report
Object Detection COCO 2017 val ViDT Swin-tiny AP50 64.5 #27 of 33 Archive leaderboard report
Object Detection COCO 2017 val ViDT Swin-tiny AP75 48.7 #27 of 33 Archive leaderboard report
Object Detection COCO 2017 val ViDT Swin-tiny APL 62.1 #27 of 33 Archive leaderboard report
Object Detection COCO 2017 val ViDT Swin-tiny APM 47.6 #27 of 33 Archive leaderboard report
Object Detection COCO 2017 val ViDT Swin-tiny APS 25.9 #27 of 33 Archive leaderboard report
Object Detection COCO 2017 val ViDT Swin-tiny Param. 38M #27 of 33 Archive leaderboard report
Object Detection COCO 2017 val ViDT Swin-nano AP 40.4 #31 of 33 Archive leaderboard report
Object Detection COCO 2017 val ViDT Swin-nano AP50 59.6 #31 of 33 Archive leaderboard report
Object Detection COCO 2017 val ViDT Swin-nano AP75 43.3 #31 of 33 Archive leaderboard report
Object Detection COCO 2017 val ViDT Swin-nano APL 55.8 #31 of 33 Archive leaderboard report
Object Detection COCO 2017 val ViDT Swin-nano APM 42.5 #31 of 33 Archive leaderboard report
Object Detection COCO 2017 val ViDT Swin-nano APS 23.2 #31 of 33 Archive leaderboard report
Object Detection COCO 2017 val ViDT Swin-nano Param. 16M #31 of 33 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Absolute Position EncodingsAdamAttentionBPEDense ConnectionsDropoutLabel SmoothingLayer NormalizationLinear LayerMulti-Head AttentionPosition-Wise Feed-Forward LayerResidual ConnectionSoftmaxStochastic DepthSwin TransformerTransformer

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections