Papers › Seeing What You Miss: Vision-Language Pre-training with Semantic Completion Learning

Seeing What You Miss: Vision-Language Pre-training with Semantic Completion Learning

24 Nov 2022CVPR 2023 1arXiv:2211.13437archive 2025-07-28

Yatai Ji, RongCheng Tu, Jie Jiang, Weijie Kong, Chengfei Cai, Wenzhe Zhao, Hongfa Wang, Yujiu Yang, Wei Liu

Cross-modal alignment is essential for vision-language pre-training (VLP) models to learn the correct corresponding information across different modalities. For this purpose, inspired by the success of masked language modeling (MLM) tasks in the NLP pre-training area, numerous masked modeling tasks have been proposed for VLP to further promote cross-modal interactions. The core idea of previous masked modeling tasks is to focus on reconstructing the masked tokens based on visible context for learning local-to-local alignment. However, most of them pay little attention to the global semantic features generated for the masked data, resulting in a limited cross-modal alignment ability of global representations. Therefore, in this paper, we propose a novel Semantic Completion Learning (SCL) task, complementary to existing masked modeling tasks, to facilitate global-to-local alignment. Specifically, the SCL task complements the missing semantics of masked data by capturing the corresponding information from the other modality, promoting learning more representative global features which have a great impact on the performance of downstream tasks. Moreover, we present a flexible vision encoder, which enables our model to perform image-text and video-text multimodal tasks simultaneously. Experimental results show that our proposed method obtains state-of-the-art performance on various vision-language benchmarks, such as visual question answering, image-text retrieval, and video-text retrieval.

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

iigroup/scl officialmentioned in papermentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Image-text RetrievalLanguage ModelingLanguage ModellingMasked Language ModelingQuestion AnsweringRetrievalText RetrievalVideo-Text RetrievalVisual Question AnsweringVisual Question Answering (VQA)Zero-Shot Video Retrievalcross-modal alignment

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Zero-Shot Video Retrieval LSMDC Yatai Ji et. al. text-to-video R@1 17.2 #10 of 16 Archive leaderboard report
Zero-Shot Video Retrieval LSMDC Yatai Ji et. al. text-to-video R@10 39.1 #10 of 16 Archive leaderboard report
Zero-Shot Video Retrieval LSMDC Yatai Ji et. al. text-to-video R@5 32.4 #10 of 16 Archive leaderboard report
Zero-Shot Video Retrieval MSR-VTT Yatai Ji et. al. text-to-video R@1 30.9 #22 of 41 Archive leaderboard report
Zero-Shot Video Retrieval MSR-VTT Yatai Ji et. al. text-to-video R@10 65.0 #22 of 41 Archive leaderboard report
Zero-Shot Video Retrieval MSR-VTT Yatai Ji et. al. text-to-video R@5 54.4 #22 of 41 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections