Papers › Cross-Modal Adaptive Dual Association for Text-to-Image Person Retrieval

Cross-Modal Adaptive Dual Association for Text-to-Image Person Retrieval

4 Dec 2023arXiv:2312.01745archive 2025-07-28

Dixuan Lin, Yixing Peng, Jingke Meng, Wei-Shi Zheng

Text-to-image person re-identification (ReID) aims to retrieve images of a person based on a given textual description. The key challenge is to learn the relations between detailed information from visual and textual modalities. Existing works focus on learning a latent space to narrow the modality gap and further build local correspondences between two modalities. However, these methods assume that image-to-text and text-to-image associations are modality-agnostic, resulting in suboptimal associations. In this work, we show the discrepancy between image-to-text association and text-to-image association and propose CADA: Cross-Modal Adaptive Dual Association that finely builds bidirectional image-text detailed associations. Our approach features a decoder-based adaptive dual association module that enables full interaction between visual and textual modalities, allowing for bidirectional and adaptive cross-modal correspondence associations. Specifically, the paper proposes a bidirectional association mechanism: Association of text Tokens to image Patches (ATP) and Association of image Regions to text Attributes (ARA). We adaptively model the ATP based on the fact that aggregating cross-modal features based on mistaken associations will lead to feature distortion. For modeling the ARA, since the attributes are typically the first distinguishing cues of a person, we propose to explore the attribute-level association by predicting the masked text phrase using the related image region. Finally, we learn the dual associations between texts and images, and the experimental results demonstrate the superiority of our dual formulation. Codes will be made publicly available.

PaperPDF

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

AttributeCross-Modal Person Re-IdentificationImage to textPerson Re-IdentificationPerson RetrievalRetrievalText based Person RetrievalText-based Person Retrieval

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Text based Person Retrieval CUHK-PEDES CADA Rank-1 78.37 #21 of 21 Archive leaderboard report
Text based Person Retrieval CUHK-PEDES CADA Rank-10 94.58 #21 of 21 Archive leaderboard report
Text based Person Retrieval CUHK-PEDES CADA Rank-5 91.57 #21 of 21 Archive leaderboard report
Text based Person Retrieval CUHK-PEDES CADA mAP 68.87 #21 of 21 Archive leaderboard report
Text based Person Retrieval ICFG-PEDES CADA Rank-1 67.81 #12 of 12 Archive leaderboard report
Text based Person Retrieval ICFG-PEDES CADA Rank-10 87.14 #12 of 12 Archive leaderboard report
Text based Person Retrieval ICFG-PEDES CADA Rank-5 82.34 #12 of 12 Archive leaderboard report
Text based Person Retrieval ICFG-PEDES CADA mAP 39.85 #12 of 12 Archive leaderboard report
Text based Person Retrieval RSTPReid CADA Rank-1 69.6 #9 of 9 Archive leaderboard report
Text based Person Retrieval RSTPReid CADA Rank-10 92.4 #9 of 9 Archive leaderboard report
Text based Person Retrieval RSTPReid CADA Rank-5 86.75 #9 of 9 Archive leaderboard report
Text based Person Retrieval RSTPReid CADA mAP 52.74 #9 of 9 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Focus

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections