{"url":"/method/dino","slug":"dino","name":"DINO","full_name":"self-DIstillation with NO labels","full_name_withheld":false,"description_markdown":"**DINO** (self-distillation with no labels) is a self-supervised learning method that directly predicts the output of a teacher network - built with a momentum encoder - using a standard cross-entropy loss. \r\n\r\nIn the example to the right, DINO is illustrated in the case of one single pair of views $\\left(x\\_{1}, x\\_{2}\\right)$ for simplicity.\r\nThe model passes two different random transformations of an input image to the student and teacher networks. Both networks have the same architecture but other parameters.\r\nThe output of the teacher network is centered with a mean computed over the batch. Each network outputs a $K$ dimensional feature normalized with a temperature [softmax](https://paperswithcode.com/method/softmax) over the feature dimension.\r\nTheir similarity is then measured with a cross-entropy loss.\r\nA stop-gradient (sg) operator is applied to the teacher to propagate gradients only through the student.\r\nThe teacher parameters are updated with the student parameters' exponential moving average (ema).","description_state":"present","introduced_year":null,"introduced_by":{"title":"Emerging Properties in Self-Supervised Vision Transformers","paper":"/paper/emerging-properties-in-self-supervised-vision","first_author":"Mathilde Caron","n_authors":7,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/emerging-properties-in-self-supervised-vision"},"source":{"url":"https://arxiv.org/abs/2104.14294v2","title":"Emerging Properties in Self-Supervised Vision Transformers","url_on_a_paper_host":true},"code_snippet_url":"https://github.com/facebookresearch/dino/blob/main/main_dino.py","code_snippet_url_on_a_code_host":true,"categories":[{"area":"Computer Vision","area_id":"computer-vision","collection":"Vision Transformers","url":"/methods/category/vision-transformers","pwc_aliases":["vision-transformer"]},{"area":"General","area_id":"general","collection":"Self-Supervised Learning","url":"/methods/category/self-supervised-learning","pwc_aliases":[]}],"n_papers_tagged":208,"archive_num_papers":208,"papers_newest_first":[{"paper":"/paper/feed-forward-scenedino-for-unsupervised","title":"Feed-Forward SceneDINO for Unsupervised Semantic Scene Completion","date":"2025-07-08","arxiv_id":"2507.06230","n_code_links":1,"syntology":null},{"paper":null,"title":"GroundingDINO-US-SAM: Text-Prompted Multi-Organ Segmentation in Ultrasound with LoRA-Tuned Vision-Language Models","date":"2025-06-30","arxiv_id":"2506.23903","n_code_links":0,"syntology":null},{"paper":null,"title":"Rethinking Random Masking in Self Distillation on ViT","date":"2025-06-12","arxiv_id":"2506.10582","n_code_links":0,"syntology":null},{"paper":null,"title":"Textile Analysis for Recycling Automation using Transfer Learning and Zero-Shot Foundation Models","date":"2025-06-06","arxiv_id":"2506.06569","n_code_links":0,"syntology":null},{"paper":"/paper/attacking-attention-of-foundation-models","title":"Attacking Attention of Foundation Models Disrupts Downstream Tasks","date":"2025-06-03","arxiv_id":"2506.05394","n_code_links":1,"syntology":null},{"paper":null,"title":"Talk2SAM: Text-Guided Semantic Enhancement for Complex-Shaped Object Segmentation","date":"2025-06-03","arxiv_id":"2506.05396","n_code_links":0,"syntology":null},{"paper":null,"title":"DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models","date":"2025-05-29","arxiv_id":"2505.24025","n_code_links":0,"syntology":null},{"paper":null,"title":"UP-SLAM: Adaptively Structured Gaussian SLAM with Uncertainty Prediction in Dynamic Environments","date":"2025-05-28","arxiv_id":"2505.22335","n_code_links":0,"syntology":null},{"paper":null,"title":"Regularized Personalization of Text-to-Image Diffusion Models without Distributional Drift","date":"2025-05-26","arxiv_id":"2505.19519","n_code_links":0,"syntology":null},{"paper":null,"title":"Unleashing Diffusion Transformers for Visual Correspondence by Modulating Massive Activations","date":"2025-05-24","arxiv_id":"2505.18584","n_code_links":0,"syntology":null},{"paper":"/paper/ssps-self-supervised-positive-sampling-for","title":"SSPS: Self-Supervised Positive Sampling for Robust Self-Supervised Speaker Verification","date":"2025-05-20","arxiv_id":"2505.14561","n_code_links":1,"syntology":null},{"paper":null,"title":"Guiding Diffusion with Deep Geometric Moments: Balancing Fidelity and Variation","date":"2025-05-18","arxiv_id":"2505.12486","n_code_links":0,"syntology":null},{"paper":null,"title":"IMAGE-ALCHEMY: Advancing subject fidelity in personalised text-to-image generation","date":"2025-05-15","arxiv_id":"2505.10743","n_code_links":0,"syntology":null},{"paper":null,"title":"BridgeIV: Bridging Customized Image and Video Generation through Test-Time Autoregressive Identity Propagation","date":"2025-05-11","arxiv_id":"2505.06985","n_code_links":0,"syntology":null},{"paper":"/paper/univla-learning-to-act-anywhere-with-task","title":"UniVLA: Learning to Act Anywhere with Task-centric Latent Actions","date":"2025-05-09","arxiv_id":"2505.06111","n_code_links":1,"syntology":{"ran":1,"of":5,"unverified":4,"pointer_only":0}},{"paper":"/paper/declip-decoupled-learning-for-open-vocabulary","title":"DeCLIP: Decoupled Learning for Open-Vocabulary Dense Perception","date":"2025-05-07","arxiv_id":"2505.04410","n_code_links":1,"syntology":{"ran":0,"of":1,"unverified":1,"pointer_only":0}},{"paper":null,"title":"From Word to Sentence: A Large-Scale Multi-Instance Dataset for Open-Set Aerial Detection","date":"2025-05-06","arxiv_id":"2505.03334","n_code_links":0,"syntology":null},{"paper":null,"title":"Pixel3DMM: Versatile Screen-Space Priors for Single-Image 3D Face Reconstruction","date":"2025-05-01","arxiv_id":"2505.00615","n_code_links":0,"syntology":null},{"paper":null,"title":"Automated Measurement of Eczema Severity with Self-Supervised Learning","date":"2025-04-21","arxiv_id":"2504.15193","n_code_links":0,"syntology":null},{"paper":"/paper/prompthmr-promptable-human-mesh-recovery","title":"PromptHMR: Promptable Human Mesh Recovery","date":"2025-04-08","arxiv_id":"2504.06397","n_code_links":1,"syntology":null},{"paper":null,"title":"Resilience of Vision Transformers for Domain Generalisation in the Presence of Out-of-Distribution Noisy Images","date":"2025-04-05","arxiv_id":"2504.04225","n_code_links":0,"syntology":null},{"paper":null,"title":"AC-LoRA: Auto Component LoRA for Personalized Artistic Style Image Generation","date":"2025-04-03","arxiv_id":"2504.02231","n_code_links":0,"syntology":null},{"paper":"/paper/f-vita-foundation-model-guided-visible-to","title":"F-ViTA: Foundation Model Guided Visible to Thermal Translation","date":"2025-04-03","arxiv_id":"2504.02801","n_code_links":1,"syntology":null},{"paper":null,"title":"Efficient Adaptation For Remote Sensing Visual Grounding","date":"2025-03-29","arxiv_id":"2503.23083","n_code_links":0,"syntology":null},{"paper":"/paper/large-self-supervised-models-bridge-the-gap","title":"Large Self-Supervised Models Bridge the Gap in Domain Adaptive Object Detection","date":"2025-03-29","arxiv_id":"2503.23220","n_code_links":1,"syntology":null},{"paper":"/paper/z-saslm-zero-shot-style-aligned-sli-blending-1","title":"Z-SASLM: Zero-Shot Style-Aligned SLI Blending Latent Manipulation","date":"2025-03-29","arxiv_id":"2503.23234","n_code_links":1,"syntology":null},{"paper":"/paper/surg-3m-a-dataset-and-foundation-model-for","title":"Surg-3M: A Dataset and Foundation Model for Perception in Surgical Settings","date":"2025-03-25","arxiv_id":"2503.19740","n_code_links":1,"syntology":null},{"paper":null,"title":"Text-Guided Image Invariant Feature Learning for Robust Image Watermarking","date":"2025-03-18","arxiv_id":"2503.13805","n_code_links":0,"syntology":null},{"paper":null,"title":"CleverDistiller: Simple and Spatially Consistent Cross-modal Distillation","date":"2025-03-12","arxiv_id":"2503.09878","n_code_links":0,"syntology":null},{"paper":null,"title":"Object-Aware DINO (Oh-A-Dino): Enhancing Self-Supervised Representations for Multi-Object Instance Retrieval","date":"2025-03-12","arxiv_id":"2503.09867","n_code_links":0,"syntology":null}],"papers_shown":30,"tasks":[{"task":"/task/semantic-segmentation","name":"Semantic Segmentation","papers":44},{"task":"/task/object-detection","name":"Object Detection","papers":38},{"task":"/task/object-detection-1","name":"object-detection","papers":33},{"task":"/task/segmentation","name":"Segmentation","papers":29},{"task":"/task/self-supervised-learning","name":"Self-Supervised Learning","papers":28},{"task":"/task/object","name":"Object","papers":25},{"task":"/task/contrastive-learning","name":"Contrastive Learning","papers":13},{"task":"/task/representation-learning","name":"Representation Learning","papers":13},{"task":"/task/image-classification","name":"Image Classification","papers":11},{"task":"/task/instance-segmentation","name":"Instance Segmentation","papers":11},{"task":"/task/image-generation","name":"Image Generation","papers":10},{"task":"/task/image-classification","name":"image-classification","papers":9},{"task":"/task/clustering","name":"Clustering","papers":8},{"task":"/task/data-augmentation","name":"Data Augmentation","papers":7},{"task":"/task/scene-understanding","name":"Scene Understanding","papers":7},{"task":"/task/transfer-learning","name":"Transfer Learning","papers":7},{"task":"/task/image-segmentation","name":"Image Segmentation","papers":6},{"task":"/task/retrieval","name":"Retrieval","papers":6},{"task":"/task/unsupervised-semantic-segmentation","name":"Unsupervised Semantic Segmentation","papers":6},{"task":"/task/anomaly-detection","name":"Anomaly Detection","papers":5}],"tasks_shown":20,"n_tasks":225,"usage_by_year":[{"year":"2021","papers":1},{"year":"2022","papers":5},{"year":"2023","papers":44},{"year":"2024","papers":108},{"year":"2025","papers":50}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/dino"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}