Papers › Prismer: A Vision-Language Model with Multi-Task Experts

Prismer: A Vision-Language Model with Multi-Task Experts

4 Mar 2023arXiv:2303.02506archive 2025-07-28

Shikun Liu, Linxi Fan, Edward Johns, Zhiding Yu, Chaowei Xiao, Anima Anandkumar

Recent vision-language models have shown impressive multi-modal generation capabilities. However, typically they require training huge models on massive datasets. As a more scalable alternative, we introduce Prismer, a data- and parameter-efficient vision-language model that leverages an ensemble of task-specific experts. Prismer only requires training of a small number of components, with the majority of network weights inherited from multiple readily-available, pre-trained experts, and kept frozen during training. By leveraging experts from a wide range of domains, we show Prismer can efficiently pool this expert knowledge and adapt it to various vision-language reasoning tasks. In our experiments, we show that Prismer achieves fine-tuned and few-shot learning performance which is competitive with current state-of-the-arts, whilst requiring up to two orders of magnitude less training data. Code is available at https://github.com/NVlabs/prismer.

PaperPDFCode

In Syntology View this paper on Syntology: its repositories, every harvested function with whether it ran, its licence and the call to fetch it.

Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

nvlabs/prismer officialmentioned in papermentioned on GitHubpytorchNOASSERTION report
KastanDay/video-pretrained-transformer mentioned on GitHubpytorchMIT report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Few-Shot LearningImage CaptioningLanguage ModelingLanguage ModellingVisual Question Answering (VQA)

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Image Captioning COCO Captions Prismer BLEU-4 40.4 #18 of 41 Archive leaderboard report
Image Captioning COCO Captions Prismer CIDER 136.5 #18 of 41 Archive leaderboard report
Image Captioning COCO Captions Prismer METEOR 31.4 #18 of 41 Archive leaderboard report
Image Captioning COCO Captions Prismer SPICE 24.4 #18 of 41 Archive leaderboard report
Image Captioning nocaps entire Prismer B1 84.87 #5 of 39 Archive leaderboard report
Image Captioning nocaps entire Prismer B2 69.99 #5 of 39 Archive leaderboard report
Image Captioning nocaps entire Prismer B3 52.48 #5 of 39 Archive leaderboard report
Image Captioning nocaps entire Prismer B4 33.66 #5 of 39 Archive leaderboard report
Image Captioning nocaps entire Prismer CIDEr 110.84 #5 of 39 Archive leaderboard report
Image Captioning nocaps entire Prismer METEOR 31.13 #5 of 39 Archive leaderboard report
Image Captioning nocaps entire Prismer ROUGE-L 60.55 #5 of 39 Archive leaderboard report
Image Captioning nocaps entire Prismer SPICE 14.91 #5 of 39 Archive leaderboard report
Image Captioning nocaps val Prismer CIDEr 107.9 #1 of 3 Archive leaderboard report
Image Captioning nocaps val Prismer SPICE 14.8 #1 of 3 Archive leaderboard report
Visual Question Answering (VQA) VQA v2 test-dev Prismer Accuracy 78.43 #16 of 56 Archive leaderboard report
Visual Question Answering (VQA) VQA v2 test-std Prismer number 61.39 #11 of 38 Archive leaderboard report
Visual Question Answering (VQA) VQA v2 test-std Prismer other 69.70 #11 of 38 Archive leaderboard report
Visual Question Answering (VQA) VQA v2 test-std Prismer overall 78.49 #11 of 38 Archive leaderboard report
Visual Question Answering (VQA) VQA v2 test-std Prismer yes/no 93.09 #11 of 38 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections