Papers › CogVLM: Visual Expert for Pretrained Language Models

CogVLM: Visual Expert for Pretrained Language Models

6 Nov 2023arXiv:2311.03079archive 2025-07-28

Weihan Wang, Qingsong Lv, Wenmeng Yu, Wenyi Hong, Ji Qi, Yan Wang, Junhui Ji, Zhuoyi Yang, Lei Zhao, Xixuan Song, Jiazheng Xu, Bin Xu, Juanzi Li, Yuxiao Dong, Ming Ding, Jie Tang

We introduce CogVLM, a powerful open-source visual language foundation model. Different from the popular shallow alignment method which maps image features into the input space of language model, CogVLM bridges the gap between the frozen pretrained language model and image encoder by a trainable visual expert module in the attention and FFN layers. As a result, CogVLM enables deep fusion of vision language features without sacrificing any performance on NLP tasks. CogVLM-17B achieves state-of-the-art performance on 10 classic cross-modal benchmarks, including NoCaps, Flicker30k captioning, RefCOCO, RefCOCO+, RefCOCOg, Visual7W, GQA, ScienceQA, VizWiz VQA and TDIUC, and ranks the 2nd on VQAv2, OKVQA, TextVQA, COCO captioning, etc., surpassing or matching PaLI-X 55B. Codes and checkpoints are available at https://github.com/THUDM/CogVLM.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

thudm/cogvlm officialmentioned in papermentioned on GitHubpytorchApache-2.0 report
THUDM/CogAgent mentioned on GitHubpytorchApache-2.0 report
MS-P3/code5 mindspore report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

1 Image, 2*2 StitchingFS-MEVQAImage RetrievalLanguage ModelingLanguage ModellingLong-Context UnderstandingVisual Question AnsweringVisual Question Answering (VQA)

1 archive task tag without a task page not shown.

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
FS-MEVQA SME GLM-4V #Learning Samples (N) 16 #5 of 7 Archive leaderboard report
FS-MEVQA SME GLM-4V ACC 34.23 #5 of 7 Archive leaderboard report
FS-MEVQA SME GLM-4V BLEU-4 14.45 #5 of 7 Archive leaderboard report
FS-MEVQA SME GLM-4V CIDEr 127.37 #5 of 7 Archive leaderboard report
FS-MEVQA SME GLM-4V Detection 0.89 #5 of 7 Archive leaderboard report
FS-MEVQA SME GLM-4V METEOR 17.53 #5 of 7 Archive leaderboard report
FS-MEVQA SME GLM-4V ROUGE-L 24.28 #5 of 7 Archive leaderboard report
FS-MEVQA SME GLM-4V SPICE 17.70 #5 of 7 Archive leaderboard report
Long-Context Understanding MMNeedle CogVLM2-Llama-3 1 Image, 2*2 Stitching, Exact Accuracy 7.3 #9 of 12 Archive leaderboard report
Long-Context Understanding MMNeedle CogVLM2-Llama-3 1 Image, 4*4 Stitching, Exact Accuracy 0.9 #9 of 12 Archive leaderboard report
Long-Context Understanding MMNeedle CogVLM2-Llama-3 1 Image, 8*8 Stitching, Exact Accuracy 0.1 #9 of 12 Archive leaderboard report
Long-Context Understanding MMNeedle CogVLM2-Llama-3 10 Images, 1*1 Stitching, Exact Accuracy 0 #9 of 12 Archive leaderboard report
Long-Context Understanding MMNeedle CogVLM2-Llama-3 10 Images, 2*2 Stitching, Exact Accuracy 0 #9 of 12 Archive leaderboard report
Long-Context Understanding MMNeedle CogVLM2-Llama-3 10 Images, 4*4 Stitching, Exact Accuracy 0 #9 of 12 Archive leaderboard report
Long-Context Understanding MMNeedle CogVLM2-Llama-3 10 Images, 8*8 Stitching, Exact Accuracy 0 #9 of 12 Archive leaderboard report
Long-Context Understanding MMNeedle CogVLM-17B 1 Image, 2*2 Stitching, Exact Accuracy 0 #11 of 12 Archive leaderboard report
Long-Context Understanding MMNeedle CogVLM-17B 1 Image, 4*4 Stitching, Exact Accuracy 0.1 #11 of 12 Archive leaderboard report
Long-Context Understanding MMNeedle CogVLM-17B 1 Image, 8*8 Stitching, Exact Accuracy 0.3 #11 of 12 Archive leaderboard report
Long-Context Understanding MMNeedle CogVLM-17B 10 Images, 1*1 Stitching, Exact Accuracy 0 #11 of 12 Archive leaderboard report
Long-Context Understanding MMNeedle CogVLM-17B 10 Images, 2*2 Stitching, Exact Accuracy 0 #11 of 12 Archive leaderboard report
Long-Context Understanding MMNeedle CogVLM-17B 10 Images, 4*4 Stitching, Exact Accuracy 0 #11 of 12 Archive leaderboard report
Long-Context Understanding MMNeedle CogVLM-17B 10 Images, 8*8 Stitching, Exact Accuracy 0 #11 of 12 Archive leaderboard report
Visual Question Answering MM-Vet GLM4 Vision GPT-4 score 63.9 #25 of 231 Archive leaderboard report
Visual Question Answering MM-Vet CogVLM(Vicuna-7B) GPT-4 score 52.8 #51 of 231 Archive leaderboard report
Visual Question Answering MM-Vet CogVLM(Vicuna-7B) Params 17B #51 of 231 Archive leaderboard report
Visual Question Answering MM-Vet v2 CogVLM-Chat GPT-4 score 45.1±0.2 #17 of 24 Archive leaderboard report
Visual Question Answering (VQA) InfiMM-Eval CogVLM-Chat Abductive 47.88 #4 of 14 Archive leaderboard report
Visual Question Answering (VQA) InfiMM-Eval CogVLM-Chat Analogical 28.75 #4 of 14 Archive leaderboard report
Visual Question Answering (VQA) InfiMM-Eval CogVLM-Chat Deductive 36.75 #4 of 14 Archive leaderboard report
Visual Question Answering (VQA) InfiMM-Eval CogVLM-Chat Overall score 37.16 #4 of 14 Archive leaderboard report
Visual Question Answering (VQA) InfiMM-Eval CogVLM-Chat Params 17B #4 of 14 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections