Papers › NAPReg: Nouns As Proxies Regularization for Semantically Aware Cross-Modal Embeddings

NAPReg: Nouns As Proxies Regularization for Semantically Aware Cross-Modal Embeddings

7 Jan 2023IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) 2023 1archive 2025-07-28

Bhavin Jawade, Deen Dayal Mohan, Naji Mohamed Ali, Srirangaraj Setlur, Venu Govindaraju

Cross-modal retrieval is a fundamental vision-language task with a broad range of practical applications. Text-to-image matching is the most common form of cross-modal retrieval where, given a large database of images and a textual query, the task is to retrieve the most relevant set of images. Existing methods utilize dual encoders with an attention mechanism and a ranking loss for learning embeddings that can be used for retrieval based on cosine similarity. Despite the fact that these methods attempt to perform semantic alignment across visual regions and textual words using tailored attention mechanisms, there is no explicit supervision from the training objective to enforce such alignment. To address this, we propose NAPReg, a novel regularization formulation that projects high-level semantic entities i.e Nouns into the embedding space as shared learnable proxies. We show that using such a formulation allows the attention mechanism to learn better word-region alignment while also utilizing region information from other samples to build a more generalized latent representation for semantic concepts. Experiments on three benchmark datasets i.e. MS-COCO, Flickr30k and Flickr8k demonstrate that our method achieves state-of-the-art results in cross-modal metric learning for text-image and image-text retrieval tasks. Code: https://github.com/bhavinjawade/NAPReq

PaperPDFCode

Code

bhavinjawade/NAPReq mentioned in paper report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Cross-Modal RetrievalImage-text RetrievalMetric LearningRetrievalText Retrieval

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Cross-Modal Retrieval COCO 2014 NAPReg Image-to-text R@1 59.8 #26 of 36 Archive leaderboard report
Cross-Modal Retrieval COCO 2014 NAPReg Text-to-image R@1 43.0 #26 of 36 Archive leaderboard report
Cross-Modal Retrieval Flickr-8k NAPReg Image-to-text R@1 56.2 #1 of 1 Archive leaderboard report
Cross-Modal Retrieval Flickr-8k NAPReg Text-to-image R@1 39.2 #1 of 1 Archive leaderboard report
Cross-Modal Retrieval Flickr30k NAPReg Image-to-text R@1 79.6 #16 of 27 Archive leaderboard report
Cross-Modal Retrieval Flickr30k NAPReg Text-to-image R@1 60.0 #16 of 27 Archive leaderboard report
Cross-Modal Retrieval MS-COCO-2014 NAPReg Text-to-image R@1 43.0 #1 of 1 Archive leaderboard report
Cross-Modal Retrieval MSCOCO-1k NAPReg Image-to-text R@1 81.9 #1 of 2 Archive leaderboard report
Cross-Modal Retrieval MSCOCO-1k NAPReg Text-to-image R@1 66.9 #1 of 2 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections