Papers › Could Giant Pretrained Image Models Extract Universal Representations?

Could Giant Pretrained Image Models Extract Universal Representations?

3 Nov 2022arXiv:2211.02043archive 2025-07-28

Yutong Lin, Ze Liu, Zheng Zhang, Han Hu, Nanning Zheng, Stephen Lin, Yue Cao

Frozen pretrained models have become a viable alternative to the pretraining-then-finetuning paradigm for transfer learning. However, with frozen models there are relatively few parameters available for adapting to downstream tasks, which is problematic in computer vision where tasks vary significantly in input/output format and the type of information that is of value. In this paper, we present a study of frozen pretrained models when applied to diverse and representative computer vision tasks, including object detection, semantic segmentation and video action recognition. From this empirical analysis, our work answers the questions of what pretraining task fits best with this frozen setting, how to make the frozen setting more flexible to various downstream tasks, and the effect of larger model sizes. We additionally examine the upper bound of performance using a giant frozen pretrained model with 3 billion parameters (SwinV2-G) and find that it reaches competitive performance on a varied set of major benchmarks with only one shared frozen base network: 60.0 box mAP and 52.2 mask mAP on COCO object detection test-dev, 57.6 val mIoU on ADE20K semantic segmentation, and 81.7 top-1 accuracy on Kinetics-400 action recognition. With this work, we hope to bring greater attention to this promising path of freezing pretrained image models.

PaperPDF

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Action RecognitionAction Recognition In VideosInstance SegmentationObject DetectionSemantic SegmentationTemporal Action LocalizationTransfer Learningobject-detection

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Action Recognition In Videos Kinetics-400 Frozen Backbone, SwinV2-G-ext22K (Video-Swin) Top-1 Accuracy 81.7 #3 of 3 Archive leaderboard report
Instance Segmentation COCO minival Frozen Backbone, SwinV2-G-ext22K (HTC) mask AP 51.6 #19 of 93 Archive leaderboard report
Object Detection COCO minival Frozen Backbone, SwinV2-G-ext22K (HTC) box AP 59.3 #29 of 220 Archive leaderboard report
Semantic Segmentation ADE20K Frozen Backbone, SwinV2-G-ext22K (Mask2Former) Validation mIoU 57.6 #31 of 235 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

BASE

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections