{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/improving-deep-visual-representation-for","title":"Improving Deep Visual Representation for Person Re-identification by Global and Local Image-language Association","arxiv_id":"1808.01571","date":"2018-08-05","proceeding":"ECCV 2018 9","authors":["Dapeng Chen","Hongsheng Li","Xihui Liu","Yantao Shen","Zejian yuan","Xiaogang Wang"],"abstract":"Person re-identification is an important task that requires learning\ndiscriminative visual features for distinguishing different person identities.\nDiverse auxiliary information has been utilized to improve the visual feature\nlearning. In this paper, we propose to exploit natural language description as\nadditional training supervisions for effective visual features. Compared with\nother auxiliary information, language can describe a specific person from more\ncompact and semantic visual aspects, thus is complementary to the pixel-level\nimage data. Our method not only learns better global visual feature with the\nsupervision of the overall description but also enforces semantic consistencies\nbetween local visual and linguistic features, which is achieved by building\nglobal and local image-language associations. The global image-language\nassociation is established according to the identity labels, while the local\nassociation is based upon the implicit correspondences between image regions\nand noun phrases. Extensive experiments demonstrate the effectiveness of\nemploying language as training supervisions with the two association schemes.\nOur method achieves state-of-the-art performance without utilizing any\nauxiliary information during testing and shows better performance than other\njoint embedding methods for the image-language association.","url_abs":"http://arxiv.org/abs/1808.01571v1","url_pdf":"http://arxiv.org/pdf/1808.01571v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"person-re-identification","task_name":"Person Re-Identification"},{"task_slug":"nlp-based-person-retrival","task_name":"Text based Person Retrieval"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/nlp-based-person-retrival-on-cuhk-pedes","task":"Text based Person Retrieval","dataset":"CUHK-PEDES","model":"GLA","rank_in_archive_order":20,"of":21,"metrics":{"R@1":"43.58","R@10":"76.26","R@5":"66.93"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/1808.01571","atlas_url":"https://app.syntology.ai/?focus=1808.01571","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}