Papers › Towards All-in-one Pre-training via Maximizing Multi-modal Mutual Information

Towards All-in-one Pre-training via Maximizing Multi-modal Mutual Information

17 Nov 2022CVPR 2023 1arXiv:2211.09807archive 2025-07-28

Weijie Su, Xizhou Zhu, Chenxin Tao, Lewei Lu, Bin Li, Gao Huang, Yu Qiao, Xiaogang Wang, Jie zhou, Jifeng Dai

To effectively exploit the potential of large-scale models, various pre-training strategies supported by massive data from different sources are proposed, including supervised pre-training, weakly-supervised pre-training, and self-supervised pre-training. It has been proved that combining multiple pre-training strategies and data from various modalities/sources can greatly boost the training of large-scale models. However, current works adopt a multi-stage pre-training system, where the complex pipeline may increase the uncertainty and instability of the pre-training. It is thus desirable that these strategies can be integrated in a single-stage manner. In this paper, we first propose a general multi-modal mutual information formula as a unified optimization target and demonstrate that all existing approaches are special cases of our framework. Under this unified perspective, we propose an all-in-one single-stage pre-training approach, named Maximizing Multi-modal Mutual Information Pre-training (M3I Pre-training). Our approach achieves better performance than previous pre-training methods on various vision benchmarks, including ImageNet classification, COCO object detection, LVIS long-tailed object detection, and ADE20k semantic segmentation. Notably, we successfully pre-train a billion-level parameter image backbone and achieve state-of-the-art performance on various benchmarks. Code shall be released at https://github.com/OpenGVLab/M3I-Pretraining.

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

OpenGVLab/M3I-Pretraining officialmentioned in papermentioned on GitHub report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

AllImage ClassificationLong-tailed Object DetectionObject DetectionSemantic Segmentationobject-detection

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Image Classification ImageNet M3I Pre-training (InternImage-H) Top 1 Accuracy 89.6% #14 of 1060 Archive leaderboard report
Object Detection COCO minival M3I Pre-training (InternImage-H) box AP 65.0 #3 of 220 Archive leaderboard report
Object Detection COCO test-dev M3I Pre-training (InternImage-H) box mAP 65.4 #3 of 225 Archive leaderboard report
Object Detection LVIS v1.0 minival M3I Pre-training (InternImage-H, single-scale) box AP 65.8 #4 of 6 Archive leaderboard report
Semantic Segmentation ADE20K M3I Pre-training (InternImage-H) Params (M) 1310 #4 of 235 Archive leaderboard report
Semantic Segmentation ADE20K M3I Pre-training (InternImage-H) Validation mIoU 62.9 #4 of 235 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections