{"url":"/method/dense-prediction-transformer","slug":"dense-prediction-transformer","name":"DPT","full_name":"Dense Prediction Transformer","full_name_withheld":false,"description_markdown":"**Dense Prediction Transformers** (DPT) are a type of [vision transformer](https://paperswithcode.com/method/vision-transformer) for dense prediction tasks.\r\n\r\nThe input image is transformed into tokens (orange) either by extracting non-overlapping patches followed by a linear projection of their flattened representation (DPT-Base and DPT-Large) or by applying a [ResNet](https://paperswithcode.com/method/resnet)-50 feature extractor (DPT-Hybrid). The image embedding is augmented with a positional embedding and a patch-independent readout token (red) is added. The tokens are passed through multiple [transformer](https://paperswithcode.com/method/transformer) stages. The tokens are reassembled from different stages into an image-like representation at multiple resolutions (green). Fusion modules (purple) progressively fuse and upsample the representations to generate a fine-grained prediction.","description_state":"present","introduced_year":null,"introduced_by":{"title":"Vision Transformers for Dense Prediction","paper":"/paper/vision-transformers-for-dense-prediction","first_author":"René Ranftl","n_authors":3,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/vision-transformers-for-dense-prediction"},"source":{"url":"https://arxiv.org/abs/2103.13413v1","title":"Vision Transformers for Dense Prediction","url_on_a_paper_host":true},"code_snippet_url":"https://github.com/intel-isl/DPT/blob/f43ef9e08d70a752195028a51be5e1aff227b913/dpt/models.py#L26","code_snippet_url_on_a_code_host":true,"categories":[{"area":"Computer Vision","area_id":"computer-vision","collection":"Image Models","url":"/methods/category/image-models","pwc_aliases":[]},{"area":"Computer Vision","area_id":"computer-vision","collection":"Vision Transformers","url":"/methods/category/vision-transformers","pwc_aliases":["vision-transformer"]}],"n_papers_tagged":25,"archive_num_papers":25,"papers_newest_first":[{"paper":null,"title":"Can In-Context Reinforcement Learning Recover From Reward Poisoning Attacks?","date":"2025-06-07","arxiv_id":"2506.06891","n_code_links":0,"syntology":null},{"paper":null,"title":"Filtering Learning Histories Enhances In-Context Reinforcement Learning","date":"2025-05-21","arxiv_id":"2505.15143","n_code_links":0,"syntology":null},{"paper":null,"title":"SAR Object Detection with Self-Supervised Pretraining and Curriculum-Aware Sampling","date":"2025-04-17","arxiv_id":"2504.13310","n_code_links":0,"syntology":null},{"paper":null,"title":"Random Policy Enables In-Context Reinforcement Learning within Trust Horizons","date":"2024-10-25","arxiv_id":"2410.19982","n_code_links":0,"syntology":null},{"paper":null,"title":"Theoretical limits of descending $\\ell_0$ sparse-regression ML algorithms","date":"2024-10-10","arxiv_id":"2410.07651","n_code_links":0,"syntology":null},{"paper":null,"title":"Endogenous Crashes as Phase Transitions","date":"2024-08-12","arxiv_id":"2408.06433","n_code_links":0,"syntology":null},{"paper":null,"title":"Pretraining Decision Transformers with Reward Prediction for In-Context Multi-task Structured Bandit Learning","date":"2024-06-07","arxiv_id":"2406.05064","n_code_links":0,"syntology":null},{"paper":"/paper/developmental-pretraining-dpt-for-image","title":"Developmental Pretraining (DPT) for Image Classification Networks","date":"2023-12-01","arxiv_id":"2312.00304","n_code_links":1,"syntology":null},{"paper":null,"title":"Enhancing Diffusion Models with 3D Perspective Geometry Constraints","date":"2023-12-01","arxiv_id":"2312.00944","n_code_links":0,"syntology":null},{"paper":null,"title":"Depth-guided Free-space Segmentation for a Mobile Robot","date":"2023-11-03","arxiv_id":"2311.01966","n_code_links":0,"syntology":null},{"paper":null,"title":"The serotonergic psychedelic N,N-dipropyltryptamine alters information-processing dynamics in cortical neural circuits","date":"2023-10-31","arxiv_id":"2310.20582","n_code_links":0,"syntology":null},{"paper":null,"title":"Supervised Pretraining Can Learn In-Context Reinforcement Learning","date":"2023-06-26","arxiv_id":"2306.14892","n_code_links":0,"syntology":null},{"paper":null,"title":"High-Resolution Synthetic RGB-D Datasets for Monocular Depth Estimation","date":"2023-05-02","arxiv_id":"2305.01732","n_code_links":0,"syntology":null},{"paper":"/paper/diffusion-models-and-semi-supervised-learners-1","title":"Diffusion Models and Semi-Supervised Learners Benefit Mutually with Few Labels","date":"2023-02-21","arxiv_id":"2302.10586","n_code_links":3,"syntology":{"ran":14,"of":22,"unverified":8,"pointer_only":3}},{"paper":"/paper/denoising-and-prompt-tuning-for-multi","title":"Denoising and Prompt-Tuning for Multi-Behavior Recommendation","date":"2023-02-12","arxiv_id":"2302.05862","n_code_links":1,"syntology":null},{"paper":"/paper/dptdr-deep-prompt-tuning-for-dense-passage","title":"DPTDR: Deep Prompt Tuning for Dense Passage Retrieval","date":"2022-08-24","arxiv_id":"2208.11503","n_code_links":2,"syntology":null},{"paper":"/paper/class-aware-visual-prompt-tuning-for-vision","title":"Dual Modality Prompt Tuning for Vision-Language Pre-Trained Model","date":"2022-08-17","arxiv_id":"2208.08340","n_code_links":1,"syntology":null},{"paper":null,"title":"SSDPT: Self-Supervised Dual-Path Transformer for Anomalous Sound Detection in Machine Condition Monitoring","date":"2022-08-06","arxiv_id":"2208.03421","n_code_links":0,"syntology":null},{"paper":"/paper/prompt-tuning-for-discriminative-pre-trained-1","title":"Prompt Tuning for Discriminative Pre-trained Language Models","date":"2022-05-23","arxiv_id":"2205.11166","n_code_links":1,"syntology":{"ran":7,"of":12,"unverified":5,"pointer_only":0}},{"paper":"/paper/declaration-based-prompt-tuning-for-visual","title":"Declaration-based Prompt Tuning for Visual Question Answering","date":"2022-05-05","arxiv_id":"2205.02456","n_code_links":1,"syntology":null},{"paper":null,"title":"Integration of neural network and fuzzy logic decision making compared with bilayered neural network in the simulation of daily dew point temperature","date":"2022-02-23","arxiv_id":"2202.12256","n_code_links":0,"syntology":null},{"paper":null,"title":"Towards 3D Scene Reconstruction from Locally Scale-Aligned Monocular Video Depth","date":"2022-02-03","arxiv_id":"2202.01470","n_code_links":0,"syntology":null},{"paper":"/paper/detail-preserving-transformer-for-light-field","title":"Detail-Preserving Transformer for Light Field Image Super-Resolution","date":"2022-01-02","arxiv_id":"2201.00346","n_code_links":1,"syntology":{"ran":10,"of":21,"unverified":11,"pointer_only":21}},{"paper":"/paper/dpt-deformable-patch-based-transformer-for","title":"DPT: Deformable Patch-based Transformer for Visual Recognition","date":"2021-07-30","arxiv_id":"2107.14467","n_code_links":1,"syntology":{"ran":1,"of":1,"unverified":0,"pointer_only":0}},{"paper":"/paper/vision-transformers-for-dense-prediction","title":"Vision Transformers for Dense Prediction","date":"2021-03-24","arxiv_id":"2103.13413","n_code_links":15,"syntology":{"ran":54,"of":116,"unverified":62,"pointer_only":15}}],"papers_shown":25,"tasks":[{"task":"/task/depth-estimation","name":"Depth Estimation","papers":4},{"task":null,"name":"In-Context Reinforcement Learning","papers":4},{"task":"/task/language-modelling","name":"Language Modelling","papers":4},{"task":"/task/monocular-depth-estimation","name":"Monocular Depth Estimation","papers":4},{"task":"/task/reinforcement-learning","name":"Reinforcement Learning","papers":3},{"task":"/task/semantic-segmentation","name":"Semantic Segmentation","papers":3},{"task":"/task/reinforcement-learning-2","name":"reinforcement-learning","papers":3},{"task":"/task/classification-1","name":"Classification","papers":2},{"task":"/task/decision-making","name":"Decision Making","papers":2},{"task":"/task/image-classification","name":"Image Classification","papers":2},{"task":"/task/in-context-learning","name":"In-Context Learning","papers":2},{"task":"/task/language-modeling","name":"Language Modeling","papers":2},{"task":"/task/object-detection","name":"Object Detection","papers":2},{"task":"/task/question-answering","name":"Question Answering","papers":2},{"task":"/task/reinforcement-learning-1","name":"Reinforcement Learning (RL)","papers":2},{"task":"/task/self-supervised-learning","name":"Self-Supervised Learning","papers":2},{"task":"/task/image-classification","name":"image-classification","papers":2},{"task":"/task/object-detection-1","name":"object-detection","papers":2},{"task":"/task/3d-scene-reconstruction","name":"3D Scene Reconstruction","papers":1},{"task":"/task/autonomous-driving","name":"Autonomous Driving","papers":1}],"tasks_shown":20,"n_tasks":53,"usage_by_year":[{"year":"2021","papers":2},{"year":"2022","papers":8},{"year":"2023","papers":8},{"year":"2024","papers":4},{"year":"2025","papers":3}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/dense-prediction-transformer"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}