Papers › DINOv2: Learning Robust Visual Features without Supervision
DINOv2: Learning Robust Visual Features without Supervision
Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, Mahmoud Assran, Nicolas Ballas, Wojciech Galuba, Russell Howes, Po-Yao Huang, Shang-Wen Li, Ishan Misra, Michael Rabbat, Vasu Sharma, Gabriel Synnaeve, Hu Xu, Hervé Jegou, Julien Mairal, Patrick Labatut, Armand Joulin, Piotr Bojanowski
The recent breakthroughs in natural language processing for model pretraining on large quantities of data have opened the way for similar foundation models in computer vision. These models could greatly simplify the use of images in any system by producing all-purpose visual features, i.e., features that work across image distributions and tasks without finetuning. This work shows that existing pretraining methods, especially self-supervised methods, can produce such features if trained on enough curated data from diverse sources. We revisit existing approaches and combine different techniques to scale our pretraining in terms of data and model size. Most of the technical contributions aim at accelerating and stabilizing the training at scale. In terms of data, we propose an automatic pipeline to build a dedicated, diverse, and curated image dataset instead of uncurated data, as typically done in the self-supervised literature. In terms of models, we train a ViT model (Dosovitskiy et al., 2020) with 1B parameters and distill it into a series of smaller models that surpass the best available all-purpose features, OpenCLIP (Ilharco et al., 2021) on most of the benchmarks at image and pixel levels.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
For agents, Syntology's MCP tool lists every function and class Syntology harvested from this paper and whether it ran (how to connect): get_harvested_code_for_paper(arxiv_id="2304.07193")
Code
Syntology Ran 21 of 46 code samples harvested from 10 repositories linked to this paper; 25 have no recorded run. Of those that ran: 2 ran · honoured contract; 2 ran · our draft was wrong; 2 ran · fixture could not drive it; 15 ran with no contract checked.
By repository: community (archive-listed): 46 samples from 10 repositories, 21 ran. The run record, sample by sample. “Ran” means executed on a synthesized input, not that the code is correct or reproduces the paper.
26 repositories listed; official and paper-mentioned ones first.
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
46 samples harvested; 21 ran; 2 honoured the contract we drafted; 25 have no recorded run. Read from Syntology's graph 2026-09-24; that is when this build read the record, not when the samples ran.
Licence: 12 of the 46 samples are pointer only, meaning Syntology does not serve that copy's text. This page shows no code text for any sample; each one links to its file in the repository.
Harvested from 10 repositories linked to this paper, official or community; each sample names its own and says which. “Ran” means the sample executed on a synthesized input. It does not mean the output is correct, and nothing here reproduces the paper's results. “Honoured” and “violated” refer to a contract Syntology drafted from the code itself; “our draft was wrong” and “fixture could not drive it” are failures of Syntology's instrument, not of the code.
Each sample ends with its code_sha256, Syntology's identity for that exact code. An agent fetches the stored sample with Syntology's MCP tool get_code(code_sha256="…") (how to connect); click an identity to copy that call.
Repository labels, per sample. official repository: The archive marks this repository official for the paper. named in the paper: The archive records that the paper mentions this repository; it is not marked official. community (archive-listed): In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper. found in paper text by Syntology: Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted. community: Not in the archive's code links for this paper; a community repository Syntology harvested. Samples from a repository marked official are listed first. Licence labels name the repository's licence as recorded at harvest. “Pointer only” means Syntology does not serve that copy's text, for one of four reasons: no licence file was found; the licence was not identified; the licence is recorded as permissive but that copy's record is not marked cleared; or the licence is outside the permissive list Syntology serves text under (MIT, Apache-2.0, BSD and similar). Some licences outside that list permit redistribution, such as WTFPL, and GPL-3.0 under its conditions; they are simply not on the list. Hover a licence label for the reason. File links open the file on GitHub at the default branch, which may have changed since the harvest.
8598c6b96e3a0fbb · report
0d03b714971c0501 · report
927d2139eab04c62 · report
e9df43ed81dc2b08 · report
22e81d05c408426c · report
05d603aa1a9a31eb · report
1d18daa479e53b33 · report
d502f73859574661 · report
8912a968e28ea03d · report
5fe0498944067403 · report
11ea568c90b9f65a · report
a5b98a91fdfc68fa · report
8b5fc434aefb49a7 · report
421274bff5d2b1e9 · report
389e99f57cbc96c6 · report
ba6d096d6c5fffc9 · report
c157f5b112b3a392 · report
3c04490a215299f2 · report
1844e38d40beb412 · report
e7fcb6d9ac9deaf4 · report
764b0f3c664386a8 · report
9557b19d4ebe95ac · report
912e7d3558003565 · report
6ceb2adc39304f27 · report
165b58cc341915b6 · report
2befc0a9ae12b450 · report
64bfcff23fded249 · report
d377a1c9f26035bb · report
316832e67ccf3416 · report
5d16d4d4fc573ac7 · report
5b78a828a7e415d5 · report
707671e31e8906f9 · report
7db18bed330a0eff · report
1dc3843893b5cd44 · report
2460ef2831016887 · report
4a4314d5371e5636 · report
85f7ffc01945bb72 · report
931b30d6ce9e37cb · report
d73ef6282cb6a0dc · report
5be3610fa1fee19e · report
dfc557c31ab30db4 · report
5e5ca5dc0d79533a · report
9c2af68d5f7355a0 · report
5474362c7aad954b · report
66fcba2b19055cb7 · report
c3b15e0d7daa5901 · report
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Depth Estimation | NYU-Depth V2 | DINOv2 (ViT-g/14 frozen, w/ DPT decoder) | RMS | 0.279 | #2 of 17 | Archive leaderboard | report |
| Domain Generalization | ImageNet-C | DINOv2 (ViT-g/14, frozen model, linear eval) | Number of params | 1100M | #1 of 47 | Archive leaderboard | report |
| Domain Generalization | ImageNet-C | DINOv2 (ViT-g/14, frozen model, linear eval) | mean Corruption Error (mCE) | 28.2 | #1 of 47 | Archive leaderboard | report |
| Domain Generalization | ImageNet-C | DINOv2 (ViT-L/14, frozen model, linear eval) | Number of params | 307M | #4 of 47 | Archive leaderboard | report |
| Domain Generalization | ImageNet-C | DINOv2 (ViT-L/14, frozen model, linear eval) | mean Corruption Error (mCE) | 31.5 | #4 of 47 | Archive leaderboard | report |
| Domain Generalization | ImageNet-C | DINOv2 (ViT-B/14, frozen model, linear eval) | Number of params | 85M | #19 of 47 | Archive leaderboard | report |
| Domain Generalization | ImageNet-C | DINOv2 (ViT-B/14, frozen model, linear eval) | mean Corruption Error (mCE) | 42.7 | #19 of 47 | Archive leaderboard | report |
| Domain Generalization | ImageNet-C | DINOv2 (ViT-S/14, frozen model, linear eval) | Number of params | 21M | #31 of 47 | Archive leaderboard | report |
| Domain Generalization | ImageNet-C | DINOv2 (ViT-S/14, frozen model, linear eval) | mean Corruption Error (mCE) | 54.4 | #31 of 47 | Archive leaderboard | report |
| Fine-Grained Image Classification | Oxford-IIIT Pet Dataset | DINOv2 (ViT-g/14, frozen model, linear eval) | Accuracy | 96.7 | #3 of 15 | Archive leaderboard | report |
| Image Classification | CIFAR-10 | DINOv2 (ViT-g/14, frozen model, linear eval) | Percentage correct | 99.5 | #2 of 265 | Archive leaderboard | report |
| Image Retrieval | AmsterTime | DINOv2 distilled (ViT-L/14 frozen) | mAP | 50.0 | #1 of 5 | Archive leaderboard | report |
| Image Retrieval | AmsterTime | DINOv2 (ViT-g/14 frozen) | mAP | 46.7 | #2 of 5 | Archive leaderboard | report |
| Image Retrieval | AmsterTime | DINOv2 distilled (ViT-B/14 frozen) | mAP | 45.6 | #3 of 5 | Archive leaderboard | report |
| Image Retrieval | AmsterTime | DINOv2 distilled (ViT-S/14 frozen) | mAP | 43.5 | #4 of 5 | Archive leaderboard | report |
| Monocular Depth Estimation | KITTI Eigen split | DINOv2 (ViT-g/14 frozen, w/ DPT decoder) | Delta < 1.25 | 0.968 | #40 of 79 | Archive leaderboard | report |
| Monocular Depth Estimation | KITTI Eigen split | DINOv2 (ViT-g/14 frozen, w/ DPT decoder) | Delta < 1.25^2 | 0.997 | #40 of 79 | Archive leaderboard | report |
| Monocular Depth Estimation | KITTI Eigen split | DINOv2 (ViT-g/14 frozen, w/ DPT decoder) | Delta < 1.25^3 | 0.9993 | #40 of 79 | Archive leaderboard | report |
| Monocular Depth Estimation | KITTI Eigen split | DINOv2 (ViT-g/14 frozen, w/ DPT decoder) | RMSE | 2.1128 | #40 of 79 | Archive leaderboard | report |
| Monocular Depth Estimation | KITTI Eigen split | DINOv2 (ViT-g/14 frozen, w/ DPT decoder) | RMSE log | 0.0882 | #40 of 79 | Archive leaderboard | report |
| Monocular Depth Estimation | KITTI Eigen split | DINOv2 (ViT-g/14 frozen, w/ DPT decoder) | Sq Rel | 0.1797 | #40 of 79 | Archive leaderboard | report |
| Monocular Depth Estimation | KITTI Eigen split | DINOv2 (ViT-g/14 frozen, w/ DPT decoder) | absolute relative error | 0.0652 | #40 of 79 | Archive leaderboard | report |
| Monocular Depth Estimation | NYU-Depth V2 | DINOv2 (ViT-g/14 frozen, w/ DPT decoder) | Delta < 1.25 | 0.9497 | #40 of 85 | Archive leaderboard | report |
| Monocular Depth Estimation | NYU-Depth V2 | DINOv2 (ViT-g/14 frozen, w/ DPT decoder) | Delta < 1.25^2 | 0.996 | #40 of 85 | Archive leaderboard | report |
| Monocular Depth Estimation | NYU-Depth V2 | DINOv2 (ViT-g/14 frozen, w/ DPT decoder) | Delta < 1.25^3 | 0.9994 | #40 of 85 | Archive leaderboard | report |
| Monocular Depth Estimation | NYU-Depth V2 | DINOv2 (ViT-g/14 frozen, w/ DPT decoder) | RMSE | 0.279 | #40 of 85 | Archive leaderboard | report |
| Monocular Depth Estimation | NYU-Depth V2 | DINOv2 (ViT-g/14 frozen, w/ DPT decoder) | absolute relative error | 0.0907 | #40 of 85 | Archive leaderboard | report |
| Monocular Depth Estimation | NYU-Depth V2 | DINOv2 (ViT-g/14 frozen, w/ DPT decoder) | log 10 | 0.0371 | #40 of 85 | Archive leaderboard | report |
| Self-Supervised Image Classification | ImageNet | DINOv2 (ViT-g/14 @448) | Number of Params | 1100M | #2 of 144 | Archive leaderboard | report |
| Self-Supervised Image Classification | ImageNet | DINOv2 (ViT-g/14 @448) | Top 1 Accuracy | 86.7% | #2 of 144 | Archive leaderboard | report |
| Self-Supervised Image Classification | ImageNet | DINOv2 (ViT-g/14) | Number of Params | 1100M | #3 of 144 | Archive leaderboard | report |
| Self-Supervised Image Classification | ImageNet | DINOv2 (ViT-g/14) | Top 1 Accuracy | 86.5% | #3 of 144 | Archive leaderboard | report |
| Self-Supervised Image Classification | ImageNet | DINOv2 distilled (ViT-L/14) | Number of Params | 307M | #4 of 144 | Archive leaderboard | report |
| Self-Supervised Image Classification | ImageNet | DINOv2 distilled (ViT-L/14) | Top 1 Accuracy | 86.3% | #4 of 144 | Archive leaderboard | report |
| Self-Supervised Image Classification | ImageNet | DINOv2 distilled (ViT-B/14) | Number of Params | 85M | #7 of 144 | Archive leaderboard | report |
| Self-Supervised Image Classification | ImageNet | DINOv2 distilled (ViT-B/14) | Top 1 Accuracy | 84.5% | #7 of 144 | Archive leaderboard | report |
| Self-Supervised Image Classification | ImageNet | DINOv2 distilled (ViT-S/14) | Number of Params | 21M | #17 of 144 | Archive leaderboard | report |
| Self-Supervised Image Classification | ImageNet | DINOv2 distilled (ViT-S/14) | Top 1 Accuracy | 81.1% | #17 of 144 | Archive leaderboard | report |
| Self-Supervised Image Classification | ImageNet (finetuned) | DINOv2 (ViT-g/14, 448) | Number of Params | 1100M | #1 of 65 | Archive leaderboard | report |
| Self-Supervised Image Classification | ImageNet (finetuned) | DINOv2 (ViT-g/14, 448) | Top 1 Accuracy | 88.9% | #1 of 65 | Archive leaderboard | report |
| Self-Supervised Image Classification | ImageNet (finetuned) | DINOv2 (ViT-g/14) | Number of Params | 1100M | #3 of 65 | Archive leaderboard | report |
| Self-Supervised Image Classification | ImageNet (finetuned) | DINOv2 (ViT-g/14) | Top 1 Accuracy | 88.5% | #3 of 65 | Archive leaderboard | report |
| Semantic Segmentation | ADE20K | DINOv2 (ViT-g/14 frozen model, w/ ViT-Adapter + Mask2former) | Params (M) | 1080 | #13 of 235 | Archive leaderboard | report |
| Semantic Segmentation | ADE20K | DINOv2 (ViT-g/14 frozen model, w/ ViT-Adapter + Mask2former) | Validation mIoU | 60.2 | #13 of 235 | Archive leaderboard | report |
| Semantic Segmentation | Fine-Grained Grass Segmentation Dataset | DINOv2 | mIoU | 47.57 | #8 of 10 | Archive leaderboard | report |
| Visual Place Recognition | 17 Places | DINOv2 | Recall@1 | 61.82 | #5 of 8 | Archive leaderboard | report |
| Visual Place Recognition | Baidu Mall | DINOv2 | Recall@1 | 49.21 | #6 of 8 | Archive leaderboard | report |
| Visual Place Recognition | Gardens Point | DINOv2 | Recall@1 | 71.50 | #5 of 7 | Archive leaderboard | report |
| Visual Place Recognition | Hawkins | DINOv2 | Recall@1 | 27.97 | #6 of 7 | Archive leaderboard | report |
| Visual Place Recognition | Laurel Caverns | DINOv2 | Recall@1 | 40.18 | #3 of 7 | Archive leaderboard | report |
| Visual Place Recognition | Mid-Atlantic Ridge | DINOv2 | Recall@1 | 24.75 | #6 of 7 | Archive leaderboard | report |
| Visual Place Recognition | Nardo-Air | DINOv2 | Recall@1 | 73.24 | #2 of 7 | Archive leaderboard | report |
| Visual Place Recognition | Nardo-Air R | DINOv2 | Recall@1 | 71.83 | #6 of 8 | Archive leaderboard | report |
| Visual Place Recognition | Oxford RobotCar Dataset | DINOv2 | Recall@1 | 39.79 | #5 of 7 | Archive leaderboard | report |
| Visual Place Recognition | Pittsburgh-30k-test | DINOv2 | Recall@1 | 78.32 | #20 of 22 | Archive leaderboard | report |
| Visual Place Recognition | St Lucia | DINOv2 | Recall@1 | 78.62 | #10 of 14 | Archive leaderboard | report |
| Visual Place Recognition | VP-Air | DINOv2 | Recall@1 | 45.23 | #2 of 7 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Methods
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections