{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/understanding-infographics-through-textual","title":"Understanding Infographics through Textual and Visual Tag Prediction","arxiv_id":"1709.09215","date":"2017-09-26","proceeding":null,"authors":["Zoya Bylinskii","Sami Alsheikh","Spandan Madan","Adria Recasens","Kimberli Zhong","Hanspeter Pfister","Fredo Durand","Aude Oliva"],"abstract":"We introduce the problem of visual hashtag discovery for infographics:\nextracting visual elements from an infographic that are diagnostic of its\ntopic. Given an infographic as input, our computational approach automatically\noutputs textual and visual elements predicted to be representative of the\ninfographic content. Concretely, from a curated dataset of 29K large\ninfographic images sampled across 26 categories and 391 tags, we present an\nautomated two step approach. First, we extract the text from an infographic and\nuse it to predict text tags indicative of the infographic content. And second,\nwe use these predicted text tags as a supervisory signal to localize the most\ndiagnostic visual elements from within the infographic i.e. visual hashtags. We\nreport performances on a categorization and multi-label tag prediction problem\nand compare our proposed visual hashtags to human annotations.","url_abs":"http://arxiv.org/abs/1709.09215v1","url_pdf":"http://arxiv.org/pdf/1709.09215v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"understanding-infographics-through-textual","repo_url":"https://github.com/cvzoya/visuallydata","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"diagnostic","task_name":"Diagnostic"},{"task_slug":"prediction","task_name":"Prediction"},{"task_slug":"tag","task_name":"TAG"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}