Papers › COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training

2 Dec 2024CVPR 2025 1arXiv:2412.01814archive 2025-07-28

Sanghwan Kim, Rui Xiao, Mariana-Iuliana Georgescu, Stephan Alaniz, Zeynep Akata

Vision-Language Models (VLMs) trained with contrastive loss have achieved significant advancements in various vision and language tasks. However, the global nature of the contrastive loss makes VLMs focus predominantly on foreground objects, neglecting other crucial information in the image, which limits their effectiveness in downstream tasks. To address these challenges, we propose COSMOS: CrOSs-MOdality Self-distillation for vision-language pre-training that integrates a novel text-cropping strategy and cross-attention module into a self-supervised learning framework. We create global and local views of images and texts (i.e., multi-modal augmentations), which are essential for self-distillation in VLMs. We further introduce a cross-attention module, enabling COSMOS to learn comprehensive cross-modal representations optimized via a cross-modality self-distillation loss. COSMOS consistently outperforms previous strong baselines on various zero-shot downstream tasks, including retrieval, classification, and semantic segmentation. Additionally, it surpasses CLIP-based models trained on larger datasets in visual perception and contextual understanding tasks. Code is available at https://github.com/ExplainableML/cosmos.

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

ExplainableML/cosmos officialmentioned in papermentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Self-Supervised LearningSemantic SegmentationUnsupervised Semantic Segmentation with Language-image Pre-trainingZero Shot SegmentationZero-Shot Cross-Modal Retrieval

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Unsupervised Semantic Segmentation with Language-image Pre-training ADE20K COSMOS ViT-B/16 Mean IoU (val) 17.7 #5 of 13 Archive leaderboard report
Unsupervised Semantic Segmentation with Language-image Pre-training COCO-Object COSMOS ViT-B/16 mIoU 31.3 #8 of 12 Archive leaderboard report
Unsupervised Semantic Segmentation with Language-image Pre-training COCO-Stuff-171 COSMOS ViT-B/16 mIoU 23.2 #7 of 12 Archive leaderboard report
Unsupervised Semantic Segmentation with Language-image Pre-training Cityscapes val COSMOS ViT-B/16 mIoU 34.7 #4 of 12 Archive leaderboard report
Unsupervised Semantic Segmentation with Language-image Pre-training PASCAL Context-59 COSMOS ViT-B/16 mIoU 33.7 #8 of 12 Archive leaderboard report
Unsupervised Semantic Segmentation with Language-image Pre-training PascalVOC-20 COSMOS ViT-B/16 mIoU 77.7 #8 of 10 Archive leaderboard report
Zero Shot Segmentation ADE20K training-free zero-shot segmentation COSMOS ViT-B/16 mIoU 17.7 #1 of 5 Archive leaderboard report
Zero-Shot Cross-Modal Retrieval COCO 2014 COSMOS ViT-B/16 Image-to-text R@1 68.0 #8 of 18 Archive leaderboard report
Zero-Shot Cross-Modal Retrieval COCO 2014 COSMOS ViT-B/16 Image-to-text R@10 92.5 #8 of 18 Archive leaderboard report
Zero-Shot Cross-Modal Retrieval COCO 2014 COSMOS ViT-B/16 Image-to-text R@5 87.8 #8 of 18 Archive leaderboard report
Zero-Shot Cross-Modal Retrieval COCO 2014 COSMOS ViT-B/16 Text-to-image R@1 52.5 #8 of 18 Archive leaderboard report
Zero-Shot Cross-Modal Retrieval COCO 2014 COSMOS ViT-B/16 Text-to-image R@10 84.9 #8 of 18 Archive leaderboard report
Zero-Shot Cross-Modal Retrieval COCO 2014 COSMOS ViT-B/16 Text-to-image R@5 77.2 #8 of 18 Archive leaderboard report
Zero-Shot Cross-Modal Retrieval COCO 2014 COSMOS ViT-B/32 Image-to-text R@1 64.3 #12 of 18 Archive leaderboard report
Zero-Shot Cross-Modal Retrieval COCO 2014 COSMOS ViT-B/32 Image-to-text R@10 92.0 #12 of 18 Archive leaderboard report
Zero-Shot Cross-Modal Retrieval COCO 2014 COSMOS ViT-B/32 Image-to-text R@5 86.5 #12 of 18 Archive leaderboard report
Zero-Shot Cross-Modal Retrieval COCO 2014 COSMOS ViT-B/32 Text-to-image R@1 48.4 #12 of 18 Archive leaderboard report
Zero-Shot Cross-Modal Retrieval COCO 2014 COSMOS ViT-B/32 Text-to-image R@10 82.6 #12 of 18 Archive leaderboard report
Zero-Shot Cross-Modal Retrieval COCO 2014 COSMOS ViT-B/32 Text-to-image R@5 74.2 #12 of 18 Archive leaderboard report
Zero-Shot Cross-Modal Retrieval Flickr30k COSMOS ViT-B/16 Image-to-text R@1 92.9 #4 of 22 Archive leaderboard report
Zero-Shot Cross-Modal Retrieval Flickr30k COSMOS ViT-B/16 Image-to-text R@10 99.9 #4 of 22 Archive leaderboard report
Zero-Shot Cross-Modal Retrieval Flickr30k COSMOS ViT-B/16 Image-to-text R@5 99.4 #4 of 22 Archive leaderboard report
Zero-Shot Cross-Modal Retrieval Flickr30k COSMOS ViT-B/16 Text-to-image R@1 80.3 #4 of 22 Archive leaderboard report
Zero-Shot Cross-Modal Retrieval Flickr30k COSMOS ViT-B/16 Text-to-image R@10 97.6 #4 of 22 Archive leaderboard report
Zero-Shot Cross-Modal Retrieval Flickr30k COSMOS ViT-B/16 Text-to-image R@5 95.3 #4 of 22 Archive leaderboard report
Zero-Shot Cross-Modal Retrieval Flickr30k COSMOS ViT-B/32 Image-to-text R@1 89.9 #11 of 22 Archive leaderboard report
Zero-Shot Cross-Modal Retrieval Flickr30k COSMOS ViT-B/32 Image-to-text R@10 99.3 #11 of 22 Archive leaderboard report
Zero-Shot Cross-Modal Retrieval Flickr30k COSMOS ViT-B/32 Image-to-text R@5 98.8 #11 of 22 Archive leaderboard report
Zero-Shot Cross-Modal Retrieval Flickr30k COSMOS ViT-B/32 Text-to-image R@1 76.1 #11 of 22 Archive leaderboard report
Zero-Shot Cross-Modal Retrieval Flickr30k COSMOS ViT-B/32 Text-to-image R@10 96.2 #11 of 22 Archive leaderboard report
Zero-Shot Cross-Modal Retrieval Flickr30k COSMOS ViT-B/32 Text-to-image R@5 92.8 #11 of 22 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Concatenated Skip ConnectionFocusSoftmax

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections