Methods › Computer Vision › Multi-Modal Methods › SyCoCa

Symmetrizing Contrastive Captioners with Attentive Masking for Multimodal Alignment

SyCoCa

2 papers tagged archive 2025-07-28

Introduced by Ziping Ma et al. in SyCoCa: Symmetrizing Contrastive Captioners with Attentive Masking for Multimodal Alignment

archive 2025-07-28 Description, source and code snippet are the archive's method entry.

Multimodal alignment between language and vision is the fundamental topic in current vision-language model research. Contrastive Captioners (CoCa), as a representative method, integrates Contrastive Language-Image Pretraining (CLIP) and Image Caption (IC) into a unified framework, resulting in impressive results. CLIP imposes a bidirectional constraints on global representation of entire images and sentences. Although IC conducts an unidirectional image-to-text generation on local representation, it lacks any constraint on local text-to-image reconstruction, which limits the ability to understand images at a fine-grained level when aligned with texts. To achieve multimodal alignment from both global and local perspectives, this paper proposes Symmetrizing Contrastive Captioners (SyCoCa), which introduces bidirectional interactions on images and texts across the global and local representation levels. Specifically, we expand a Text-Guided Masked Image Modeling (TG-MIM) head based on ITC and IC heads. The improved SyCoCa can further leverage textual cues to reconstruct contextual images and visual cues to predict textual contents. When implementing bidirectional local interactions, the local contents of images tend to be cluttered or unrelated to their textual descriptions. Thus, we employ an attentive masking strategy to select effective image patches for interaction. Extensive experiments on five vision-language tasks, including image-text retrieval, image-captioning, visual question answering, and zero-shot/finetuned image classification, validate the effectiveness of our proposed method.

PaperSource

Papers archive 2025-07-28

2 shown of 2, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.

Tasks archive 2025-07-28

18 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.

TaskPapers
GPU1
Image Captioning1
Image Classification1
Image Reconstruction1
Image to text1
Image-text Retrieval1
Language Modelling1
Question Answering1
Text Generation1
Text Retrieval1
Visual Question Answering1
Zero-Shot Cross-Modal Retrieval1
Zero-Shot Learning1
Zero-Shot Transfer Image Classification1
Zero-shot Image Retrieval1
Zero-shot Text-to-Image Retrieval1
image-classification1
zero-shot-classification1

Usage over time archive 2025-07-28

Papers per year tagged with SyCoCa: 2024 to 2024, peak 2 2 0 2024: 2 papers 2024
Papers per year the archive tags with this method, by the paper's archive date (2 dated). Bars are counts, not a trend claim.

Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).

Categories archive 2025-07-28

Multi-Modal Methods

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections