Papers › Dynamic Self-adaptive Multiscale Distillation from Pre-trained Multimodal Large Model...

Dynamic Self-adaptive Multiscale Distillation from Pre-trained Multimodal Large Model for Efficient Cross-modal Representation Learning

16 Apr 2024arXiv:2404.10838archive 2025-07-28

Zhengyang Liang, Meiyu Liang, Wei Huang, Yawen Li, Zhe Xue

In recent years, pre-trained multimodal large models have attracted widespread attention due to their outstanding performance in various multimodal applications. Nonetheless, the extensive computational resources and vast datasets required for their training present significant hurdles for deployment in environments with limited computational resources. To address this challenge, we propose a novel dynamic self-adaptive multiscale distillation from pre-trained multimodal large model for efficient cross-modal representation learning for the first time. Unlike existing distillation methods, our strategy employs a multiscale perspective, enabling the extraction structural knowledge across from the pre-trained multimodal large model. Ensuring that the student model inherits a comprehensive and nuanced understanding of the teacher knowledge. To optimize each distillation loss in a balanced and efficient manner, we propose a dynamic self-adaptive distillation loss balancer, a novel component eliminating the need for manual loss weight adjustments and dynamically balances each loss item during the distillation process. Our methodology streamlines pre-trained multimodal large models using only their output features and original image-level information, requiring minimal computational resources. This efficient approach is suited for various applications and allows the deployment of advanced multimodal technologies even in resource-limited settings. Extensive experiments has demonstrated that our method maintains high performance while significantly reducing model complexity and training costs. Moreover, our distilled student model utilizes only image-level information to achieve state-of-the-art performance on cross-modal retrieval tasks, surpassing previous methods that relied on region-level information.

PaperPDFCode

Code

chrisx599/dsmd officialmentioned in paperpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Cross-Modal RetrievalRepresentation Learning

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Cross-Modal Retrieval COCO 2014 DSMD Image-to-text R@1 48.0 #13 of 36 Archive leaderboard report
Cross-Modal Retrieval COCO 2014 DSMD Image-to-text R@10 84.5 #13 of 36 Archive leaderboard report
Cross-Modal Retrieval COCO 2014 DSMD Image-to-text R@5 75.6 #13 of 36 Archive leaderboard report
Cross-Modal Retrieval COCO 2014 DSMD Text-to-image R@1 62.1 #13 of 36 Archive leaderboard report
Cross-Modal Retrieval COCO 2014 DSMD Text-to-image R@10 92.0 #13 of 36 Archive leaderboard report
Cross-Modal Retrieval COCO 2014 DSMD Text-to-image R@5 85.9 #13 of 36 Archive leaderboard report
Cross-Modal Retrieval Flickr30k DSMD Image-to-text R@1 82.5 #14 of 27 Archive leaderboard report
Cross-Modal Retrieval Flickr30k DSMD Image-to-text R@10 97.7 #14 of 27 Archive leaderboard report
Cross-Modal Retrieval Flickr30k DSMD Image-to-text R@5 95.5 #14 of 27 Archive leaderboard report
Cross-Modal Retrieval Flickr30k DSMD Text-to-image R@1 68.4 #14 of 27 Archive leaderboard report
Cross-Modal Retrieval Flickr30k DSMD Text-to-image R@10 94.4 #14 of 27 Archive leaderboard report
Cross-Modal Retrieval Flickr30k DSMD Text-to-image R@5 90.8 #14 of 27 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections