{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/vitaa-visual-textual-attributes-alignment-in","title":"ViTAA: Visual-Textual Attributes Alignment in Person Search by Natural Language","arxiv_id":"2005.07327","date":"2020-05-15","proceeding":"ECCV 2020 8","authors":["Zhe Wang","Zhiyuan Fang","Jun Wang","Yezhou Yang"],"abstract":"Person search by natural language aims at retrieving a specific person in a large-scale image pool that matches the given textual descriptions. While most of the current methods treat the task as a holistic visual and textual feature matching one, we approach it from an attribute-aligning perspective that allows grounding specific attribute phrases to the corresponding visual regions. We achieve success as well as the performance boosting by a robust feature learning that the referred identity can be accurately bundled by multiple attribute visual cues. To be concrete, our Visual-Textual Attribute Alignment model (dubbed as ViTAA) learns to disentangle the feature space of a person into subspaces corresponding to attributes using a light auxiliary attribute segmentation computing branch. It then aligns these visual features with the textual attributes parsed from the sentences by using a novel contrastive learning loss. Upon that, we validate our ViTAA framework through extensive experiments on tasks of person search by natural language and by attribute-phrase queries, on which our system achieves state-of-the-art performances. Code will be publicly available upon publication.","url_abs":"https://arxiv.org/abs/2005.07327v2","url_pdf":"https://arxiv.org/pdf/2005.07327v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"vitaa-visual-textual-attributes-alignment-in","repo_url":"https://github.com/Jarr0d/ViTAA","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}},{"paper_slug":"vitaa-visual-textual-attributes-alignment-in","repo_url":"https://github.com/Jarr0d/Human-Parsing-Network","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"attribute","task_name":"Attribute"},{"task_slug":"contrastive-learning","task_name":"Contrastive Learning"},{"task_slug":"person-search","task_name":"Person Search"},{"task_slug":"nlp-based-person-retrival","task_name":"Text based Person Retrieval"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/nlp-based-person-retrival-on-cuhk-pedes","task":"Text based Person Retrieval","dataset":"CUHK-PEDES","model":"ViTAA","rank_in_archive_order":16,"of":21,"metrics":{"R@1":"55.97","R@10":"83.52","R@5":"75.84"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2005.07327","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}