Papers › DINO as a von Mises-Fisher mixture model

DINO as a von Mises-Fisher mixture model

17 May 2024ICLR 2023 2arXiv:2405.10939archive 2025-07-28

Hariprasath Govindarajan, Per Sidén, Jacob Roll, Fredrik Lindsten

Self-distillation methods using Siamese networks are popular for self-supervised pre-training. DINO is one such method based on a cross-entropy loss between K-dimensional probability vectors, obtained by applying a softmax function to the dot product between representations and learnt prototypes. Given the fact that the learned representations are L²-normalized, we show that DINO and its derivatives, such as iBOT, can be interpreted as a mixture model of von Mises-Fisher components. With this interpretation, DINO assumes equal precision for all components when the prototypes are also L²-normalized. Using this insight we propose DINO-vMF, that adds appropriate normalization constants when computing the cluster assignment probabilities. Unlike DINO, DINO-vMF is stable also for the larger ViT-Base model with unnormalized prototypes. We show that the added flexibility of the mixture model is beneficial in terms of better image representations. The DINO-vMF pre-trained model consistently performs better than DINO on a range of downstream tasks. We obtain similar improvements for iBOT-vMF vs iBOT and thereby show the relevance of our proposed modification also for other methods derived from DINO.

PaperPDFConference PDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

model

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Self-Supervised Image Classification ImageNet iBOT-vMF (ViT-B/16) Number of Params 85M #24 of 144 Archive leaderboard report
Self-Supervised Image Classification ImageNet iBOT-vMF (ViT-B/16) Top 1 Accuracy 80.3% #24 of 144 Archive leaderboard report
Self-Supervised Image Classification ImageNet DINO-vMF (ViT-B/16) Number of Params 85M #42 of 144 Archive leaderboard report
Self-Supervised Image Classification ImageNet DINO-vMF (ViT-B/16) Top 1 Accuracy 78.8% #42 of 144 Archive leaderboard report
Self-Supervised Image Classification ImageNet DINO-vMF (ViT-S/16) Number of Params 21M #57 of 144 Archive leaderboard report
Self-Supervised Image Classification ImageNet DINO-vMF (ViT-S/16) Top 1 Accuracy 77.0% #57 of 144 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

AttentionDINODense ConnectionsLayer NormalizationLinear LayerMulti-Head AttentionResidual ConnectionSoftmaxVision Transformer

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections