{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/fine-grained-visual-textual-representation","title":"Fine-grained Visual-textual Representation Learning","arxiv_id":"1709.00340","date":"2017-08-31","proceeding":null,"authors":["Xiangteng He","Yuxin Peng"],"abstract":"Fine-grained visual categorization is to recognize hundreds of subcategories\nbelonging to the same basic-level category, which is a highly challenging task\ndue to the quite subtle and local visual distinctions among similar\nsubcategories. Most existing methods generally learn part detectors to discover\ndiscriminative regions for better categorization performance. However, not all\nparts are beneficial and indispensable for visual categorization, and the\nsetting of part detector number heavily relies on prior knowledge as well as\nexperimental validation. As is known to all, when we describe the object of an\nimage via textual descriptions, we mainly focus on the pivotal characteristics,\nand rarely pay attention to common characteristics as well as the background\nareas. This is an involuntary transfer from human visual attention to textual\nattention, which leads to the fact that textual attention tells us how many and\nwhich parts are discriminative and significant to categorization. So textual\nattention could help us to discover visual attention in image. Inspired by\nthis, we propose a fine-grained visual-textual representation learning (VTRL)\napproach, and its main contributions are: (1) Fine-grained visual-textual\npattern mining devotes to discovering discriminative visual-textual pairwise\ninformation for boosting categorization performance through jointly modeling\nvision and text with generative adversarial networks (GANs), which\nautomatically and adaptively discovers discriminative parts. (2) Visual-textual\nrepresentation learning jointly combines visual and textual information, which\npreserves the intra-modality and inter-modality information to generate\ncomplementary fine-grained representation, as well as further improves\ncategorization performance.","url_abs":"http://arxiv.org/abs/1709.00340v4","url_pdf":"http://arxiv.org/pdf/1709.00340v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"fine-grained-visual-textual-representation","repo_url":"https://github.com/PKU-ICST-MIPL/OPAM_TIP2018","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"fine-grained-visual-categorization","task_name":"Fine-Grained Visual Categorization"},{"task_slug":"representation-learning","task_name":"Representation Learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1709.00340","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}