Papers › COSA: Concatenated Sample Pretrained Vision-Language Foundation Model

COSA: Concatenated Sample Pretrained Vision-Language Foundation Model

15 Jun 2023arXiv:2306.09085archive 2025-07-28

Sihan Chen, Xingjian He, Handong Li, Xiaojie Jin, Jiashi Feng, Jing Liu

Due to the limited scale and quality of video-text training corpus, most vision-language foundation models employ image-text datasets for pretraining and primarily focus on modeling visually semantic representations while disregarding temporal semantic representations and correlations. To address this issue, we propose COSA, a COncatenated SAmple pretrained vision-language foundation model. COSA jointly models visual contents and event-level temporal cues using only image-text corpora. We achieve this by sequentially concatenating multiple image-text pairs as inputs for pretraining. This transformation effectively converts existing image-text corpora into a pseudo long-form video-paragraph corpus, enabling richer scene transformations and explicit event-description correspondence. Extensive experiments demonstrate that COSA consistently improves performance across a broad range of downstream tasks, including long-form/short-form video-text tasks and image-text tasks such as retrieval, captioning, and question answering. Notably, COSA achieves state-of-the-art results on various competitive benchmarks. Code and model are released at https://github.com/TXH-mercury/COSA.

PaperPDFCode

Code

txh-mercury/cosa officialmentioned in paperpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

FormQuestion AnsweringRetrievalTGIF-FrameVideo CaptioningVideo Captioning on MSR-VTTVideo Question AnsweringVideo RetrievalVisual Question Answering (VQA)model

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Video Captioning MSR-VTT COSA BLEU-4 53.7 #5 of 24 Archive leaderboard report
Video Captioning MSR-VTT COSA CIDEr 74.7 #5 of 24 Archive leaderboard report
Video Captioning MSVD COSA BLEU-4 76.5 #4 of 16 Archive leaderboard report
Video Captioning MSVD COSA CIDEr 178.5 #4 of 16 Archive leaderboard report
Video Captioning TVC COSA BLEU-4 18.8 #2 of 2 Archive leaderboard report
Video Captioning TVC COSA CIDEr 70.7 #2 of 2 Archive leaderboard report
Video Captioning VATEX COSA BLEU-4 43.7 #3 of 10 Archive leaderboard report
Video Captioning VATEX COSA CIDEr 96.5 #3 of 10 Archive leaderboard report
Video Captioning YouCook2 COSA BLEU-4 10.1 #9 of 14 Archive leaderboard report
Video Captioning YouCook2 COSA CIDEr 1.31 #9 of 14 Archive leaderboard report
Video Question Answering ActivityNet-QA COSA Accuracy 49.9 #6 of 36 Archive leaderboard report
Video Question Answering MSRVTT-QA COSA Accuracy 49.2 #4 of 14 Archive leaderboard report
Video Retrieval ActivityNet COSA text-to-video R@1 67.3 #5 of 31 Archive leaderboard report
Video Retrieval DiDeMo COSA text-to-video R@1 70.5 #4 of 40 Archive leaderboard report
Video Retrieval LSMDC COSA text-to-video R@1 39.4 #5 of 38 Archive leaderboard report
Video Retrieval MSR-VTT COSA text-to-video R@1 57.9 #7 of 40 Archive leaderboard report
Visual Question Answering (VQA) MSVD-QA COSA Accuracy 0.60 #6 of 36 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Focus

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections