Methods › Computer Vision › Vision and Language Pre-Trained Models › WenLan
WenLan
Introduced by Yuqi Huo et al. in WenLan: Bridging Vision and Language by Large-Scale Multi-Modal Pre-Training
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
Proposes a two-tower pre-training model called BriVL within the cross-modal contrastive learning framework. A cross-modal pre-training model is defined based on the image-text retrieval task. The main goal is thus to learn two encoders that can embed image and text samples into the same space for effective image-text retrieval. To enforce such cross-modal embedding learning, we introduce contrastive learning with the InfoNCE loss into the BriVL model. Given text embedding, the learning objective aims to find the best image embedding from a batch of image embeddings. Similarly, for a given image embedding, the learning objective is to find the best text embedding from a batch of text embeddings. The pre-training model learns a cross-modal embedding space by jointly training the image and text encoders to maximize the cosine similarity of the image and text embeddings of the true pair for each sample in the batch while minimizing the cosine similarity of the embeddings of the other incorrect pairs.
Papers archive 2025-07-28
4 shown of 4, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
One Model, Multiple Modalities: A Sparsely Activated Approach for Text, Sound, Image, Video and Code 12 May 2022 · 0 repositories · arXiv:2205.06126
-
Wukong: A 100 Million Large-scale Chinese Cross-modal Pre-training Benchmark 14 Feb 2022 · 1 repository · arXiv:2202.06767Syntology ran 0 of 2 samples · 2 unverified
-
EfficientCLIP: Efficient Cross-Modal Pre-training by Ensemble Confident Learning and Language Modeling 10 Sep 2021 · 0 repositories · arXiv:2109.04699
-
WenLan: Bridging Vision and Language by Large-Scale Multi-Modal Pre-Training 11 Mar 2021 · 2 repositories · arXiv:2103.06561
Tasks archive 2025-07-28
18 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections