Papers › CogVLM: Visual Expert for Pretrained Language Models
CogVLM: Visual Expert for Pretrained Language Models
Weihan Wang, Qingsong Lv, Wenmeng Yu, Wenyi Hong, Ji Qi, Yan Wang, Junhui Ji, Zhuoyi Yang, Lei Zhao, Xixuan Song, Jiazheng Xu, Bin Xu, Juanzi Li, Yuxiao Dong, Ming Ding, Jie Tang
We introduce CogVLM, a powerful open-source visual language foundation model. Different from the popular shallow alignment method which maps image features into the input space of language model, CogVLM bridges the gap between the frozen pretrained language model and image encoder by a trainable visual expert module in the attention and FFN layers. As a result, CogVLM enables deep fusion of vision language features without sacrificing any performance on NLP tasks. CogVLM-17B achieves state-of-the-art performance on 10 classic cross-modal benchmarks, including NoCaps, Flicker30k captioning, RefCOCO, RefCOCO+, RefCOCOg, Visual7W, GQA, ScienceQA, VizWiz VQA and TDIUC, and ranks the 2nd on VQAv2, OKVQA, TextVQA, COCO captioning, etc., surpassing or matching PaLI-X 55B. Codes and checkpoints are available at https://github.com/THUDM/CogVLM.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
1 archive task tag without a task page not shown.
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| FS-MEVQA | SME | GLM-4V | #Learning Samples (N) | 16 | #5 of 7 | Archive leaderboard | report |
| FS-MEVQA | SME | GLM-4V | ACC | 34.23 | #5 of 7 | Archive leaderboard | report |
| FS-MEVQA | SME | GLM-4V | BLEU-4 | 14.45 | #5 of 7 | Archive leaderboard | report |
| FS-MEVQA | SME | GLM-4V | CIDEr | 127.37 | #5 of 7 | Archive leaderboard | report |
| FS-MEVQA | SME | GLM-4V | Detection | 0.89 | #5 of 7 | Archive leaderboard | report |
| FS-MEVQA | SME | GLM-4V | METEOR | 17.53 | #5 of 7 | Archive leaderboard | report |
| FS-MEVQA | SME | GLM-4V | ROUGE-L | 24.28 | #5 of 7 | Archive leaderboard | report |
| FS-MEVQA | SME | GLM-4V | SPICE | 17.70 | #5 of 7 | Archive leaderboard | report |
| Long-Context Understanding | MMNeedle | CogVLM2-Llama-3 | 1 Image, 2*2 Stitching, Exact Accuracy | 7.3 | #9 of 12 | Archive leaderboard | report |
| Long-Context Understanding | MMNeedle | CogVLM2-Llama-3 | 1 Image, 4*4 Stitching, Exact Accuracy | 0.9 | #9 of 12 | Archive leaderboard | report |
| Long-Context Understanding | MMNeedle | CogVLM2-Llama-3 | 1 Image, 8*8 Stitching, Exact Accuracy | 0.1 | #9 of 12 | Archive leaderboard | report |
| Long-Context Understanding | MMNeedle | CogVLM2-Llama-3 | 10 Images, 1*1 Stitching, Exact Accuracy | 0 | #9 of 12 | Archive leaderboard | report |
| Long-Context Understanding | MMNeedle | CogVLM2-Llama-3 | 10 Images, 2*2 Stitching, Exact Accuracy | 0 | #9 of 12 | Archive leaderboard | report |
| Long-Context Understanding | MMNeedle | CogVLM2-Llama-3 | 10 Images, 4*4 Stitching, Exact Accuracy | 0 | #9 of 12 | Archive leaderboard | report |
| Long-Context Understanding | MMNeedle | CogVLM2-Llama-3 | 10 Images, 8*8 Stitching, Exact Accuracy | 0 | #9 of 12 | Archive leaderboard | report |
| Long-Context Understanding | MMNeedle | CogVLM-17B | 1 Image, 2*2 Stitching, Exact Accuracy | 0 | #11 of 12 | Archive leaderboard | report |
| Long-Context Understanding | MMNeedle | CogVLM-17B | 1 Image, 4*4 Stitching, Exact Accuracy | 0.1 | #11 of 12 | Archive leaderboard | report |
| Long-Context Understanding | MMNeedle | CogVLM-17B | 1 Image, 8*8 Stitching, Exact Accuracy | 0.3 | #11 of 12 | Archive leaderboard | report |
| Long-Context Understanding | MMNeedle | CogVLM-17B | 10 Images, 1*1 Stitching, Exact Accuracy | 0 | #11 of 12 | Archive leaderboard | report |
| Long-Context Understanding | MMNeedle | CogVLM-17B | 10 Images, 2*2 Stitching, Exact Accuracy | 0 | #11 of 12 | Archive leaderboard | report |
| Long-Context Understanding | MMNeedle | CogVLM-17B | 10 Images, 4*4 Stitching, Exact Accuracy | 0 | #11 of 12 | Archive leaderboard | report |
| Long-Context Understanding | MMNeedle | CogVLM-17B | 10 Images, 8*8 Stitching, Exact Accuracy | 0 | #11 of 12 | Archive leaderboard | report |
| Visual Question Answering | MM-Vet | GLM4 Vision | GPT-4 score | 63.9 | #25 of 231 | Archive leaderboard | report |
| Visual Question Answering | MM-Vet | CogVLM(Vicuna-7B) | GPT-4 score | 52.8 | #51 of 231 | Archive leaderboard | report |
| Visual Question Answering | MM-Vet | CogVLM(Vicuna-7B) | Params | 17B | #51 of 231 | Archive leaderboard | report |
| Visual Question Answering | MM-Vet v2 | CogVLM-Chat | GPT-4 score | 45.1±0.2 | #17 of 24 | Archive leaderboard | report |
| Visual Question Answering (VQA) | InfiMM-Eval | CogVLM-Chat | Abductive | 47.88 | #4 of 14 | Archive leaderboard | report |
| Visual Question Answering (VQA) | InfiMM-Eval | CogVLM-Chat | Analogical | 28.75 | #4 of 14 | Archive leaderboard | report |
| Visual Question Answering (VQA) | InfiMM-Eval | CogVLM-Chat | Deductive | 36.75 | #4 of 14 | Archive leaderboard | report |
| Visual Question Answering (VQA) | InfiMM-Eval | CogVLM-Chat | Overall score | 37.16 | #4 of 14 | Archive leaderboard | report |
| Visual Question Answering (VQA) | InfiMM-Eval | CogVLM-Chat | Params | 17B | #4 of 14 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections