{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-novel-visual-representation-on-text-using","title":"A Novel Visual Representation on Text Using Diverse Conditional GAN for Visual Recognition","arxiv_id":null,"date":"2021-03-05","proceeding":"IEEE TIP 2021 2021 3","authors":["Tao Hu; Chengjiang Long; Chunxia Xiao"],"abstract":"Abstract— Automatic image visual recognition can make full\r\nuse of largely available images with text descriptions on social\r\nmedia platforms to build large-scale image labeled datasets.\r\nIn this paper, we propose a novel visual text representation,\r\nnamed DG-VRT (Diverse GAN-Visual Representation on Text),\r\nwhich extracts visual features from synthetic images generated by\r\na diverse conditional Generative Adversarial Network (DCGAN)\r\non the text, for visual recognition. The DCGAN incorporates\r\nthe current state-of-the-art text-to-image GANs and generates\r\nmultiple synthetic images with various prior noises conditioned\r\non a text. Then we extract deep visual features from the generated\r\nsynthetic images to explore the underlying visual concepts and\r\nprovide a visual transformation on text in feature space. Finally,\r\nwe combine image-level visual features, text-level features and\r\nvisual features based on synthetic images together to recognize\r\nthe images, and we also extend the proposed work to semantic\r\nsegmentation. We conduct extensive experiments on two benchmark datasets and the experimental results demonstrate the efficacy of our proposed representation on text for visual recognition.\r\nIndex Terms— Visual representation, diverse conditional GAN,\r\nvisual recognition.","url_abs":"https://ieeexplore.ieee.org/document/9371392","url_pdf":"https://ieeexplore.ieee.org/document/9371392","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"a-novel-visual-representation-on-text-using","repo_url":"https://github.com/WhUGraphVision/DG-VRT","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":null,"task_name":"Generative Adversarial Network"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"}],"methods":[{"method_slug":"batch-normalization","method_name":"Batch Normalization"},{"method_slug":"convolution","method_name":"Convolution"},{"method_slug":"dcgan","method_name":"DCGAN"},{"method_slug":"relu","method_name":"ReLU"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}