Methods › Computer Vision › Vision and Language Pre-Trained Models › WenLan

WenLan

4 papers tagged archive 2025-07-28

Introduced by Yuqi Huo et al. in WenLan: Bridging Vision and Language by Large-Scale Multi-Modal Pre-Training

archive 2025-07-28 Description, source and code snippet are the archive's method entry.

Proposes a two-tower pre-training model called BriVL within the cross-modal contrastive learning framework. A cross-modal pre-training model is defined based on the image-text retrieval task. The main goal is thus to learn two encoders that can embed image and text samples into the same space for effective image-text retrieval. To enforce such cross-modal embedding learning, we introduce contrastive learning with the InfoNCE loss into the BriVL model. Given text embedding, the learning objective aims to find the best image embedding from a batch of image embeddings. Similarly, for a given image embedding, the learning objective is to find the best text embedding from a batch of text embeddings. The pre-training model learns a cross-modal embedding space by jointly training the image and text encoders to maximize the cosine similarity of the image and text embeddings of the true pair for each sample in the batch while minimizing the cosine similarity of the embeddings of the other incorrect pairs.

PaperSource

Papers archive 2025-07-28

4 shown of 4, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.

Tasks archive 2025-07-28

18 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.

TaskPapers
Image Retrieval3
Retrieval3
Contrastive Learning2
Text Retrieval2
Benchmarking1
Cross-Modal Retrieval1
GPU1
Image Captioning1
Image Classification1
Image-text Retrieval1
Image-to-Text Retrieval1
Language Modeling1
Language Modelling1
Text Classification1
Zero-Shot Image Classification1
Zero-shot Image Retrieval1
image-classification1
text-classification1

Usage over time archive 2025-07-28

Papers per year tagged with WenLan: 2021 to 2022, peak 2 2 0 2021: 2 papers 2021 2022: 2 papers 2022
Papers per year the archive tags with this method, by the paper's archive date (4 dated). Bars are counts, not a trend claim.

Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).

Categories archive 2025-07-28

Vision and Language Pre-Trained Models

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections