Papers › ViSTA: Vision and Scene Text Aggregation for Cross-Modal Retrieval

ViSTA: Vision and Scene Text Aggregation for Cross-Modal Retrieval

31 Mar 2022CVPR 2022 1arXiv:2203.16778archive 2025-07-28

Mengjun Cheng, Yipeng Sun, Longchao Wang, Xiongwei Zhu, Kun Yao, Jie Chen, Guoli Song, Junyu Han, Jingtuo Liu, Errui Ding, Jingdong Wang

Visual appearance is considered to be the most important cue to understand images for cross-modal retrieval, while sometimes the scene text appearing in images can provide valuable information to understand the visual semantics. Most of existing cross-modal retrieval approaches ignore the usage of scene text information and directly adding this information may lead to performance degradation in scene text free scenarios. To address this issue, we propose a full transformer architecture to unify these cross-modal retrieval scenarios in a single Vision and Scene Text Aggregation framework (ViSTA). Specifically, ViSTA utilizes transformer blocks to directly encode image patches and fuse scene text embedding to learn an aggregated visual representation for cross-modal retrieval. To tackle the modality missing problem of scene text, we propose a novel fusion token based transformer aggregation approach to exchange the necessary scene text information only through the fusion token and concentrate on the most important features in each modality. To further strengthen the visual modality, we develop dual contrastive learning losses to embed both image-text pairs and fusion-text pairs into a common cross-modal space. Compared to existing methods, ViSTA enables to aggregate relevant scene text semantics with visual appearance, and hence improve results under both scene text free and scene text aware scenarios. Experimental results show that ViSTA outperforms other methods by at least 8.4% at Recall@1 for scene text aware retrieval task. Compared with state-of-the-art scene text free retrieval methods, ViSTA can achieve better accuracy on Flicker30K and MSCOCO while running at least three times faster during the inference stage, which validates the effectiveness of the proposed framework.

PaperPDFConference PDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Contrastive LearningCross-Modal RetrievalRetrieval

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Cross-Modal Retrieval COCO 2014 ViSTA Image-to-text R@1 68.9 #21 of 36 Archive leaderboard report
Cross-Modal Retrieval COCO 2014 ViSTA Image-to-text R@10 95.4 #21 of 36 Archive leaderboard report
Cross-Modal Retrieval COCO 2014 ViSTA Image-to-text R@5 90.1 #21 of 36 Archive leaderboard report
Cross-Modal Retrieval COCO 2014 ViSTA Text-to-image R@1 52.6 #21 of 36 Archive leaderboard report
Cross-Modal Retrieval COCO 2014 ViSTA Text-to-image R@10 87.6 #21 of 36 Archive leaderboard report
Cross-Modal Retrieval COCO 2014 ViSTA Text-to-image R@5 79.6 #21 of 36 Archive leaderboard report
Cross-Modal Retrieval Flickr30k ViSTA Image-to-text R@1 89.5 #10 of 27 Archive leaderboard report
Cross-Modal Retrieval Flickr30k ViSTA Image-to-text R@10 99.6 #10 of 27 Archive leaderboard report
Cross-Modal Retrieval Flickr30k ViSTA Image-to-text R@5 98.4 #10 of 27 Archive leaderboard report
Cross-Modal Retrieval Flickr30k ViSTA Text-to-image R@1 75.8 #10 of 27 Archive leaderboard report
Cross-Modal Retrieval Flickr30k ViSTA Text-to-image R@10 96.9 #10 of 27 Archive leaderboard report
Cross-Modal Retrieval Flickr30k ViSTA Text-to-image R@5 94.2 #10 of 27 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

AWAREContrastive Learning

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections