{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/multimodal-attribute-extraction","title":"Multimodal Attribute Extraction","arxiv_id":"1711.11118","date":"2017-11-29","proceeding":null,"authors":["Robert L. Logan IV","Samuel Humeau","Sameer Singh"],"abstract":"The broad goal of information extraction is to derive structured information\nfrom unstructured data. However, most existing methods focus solely on text,\nignoring other types of unstructured data such as images, video and audio which\ncomprise an increasing portion of the information on the web. To address this\nshortcoming, we propose the task of multimodal attribute extraction. Given a\ncollection of unstructured and semi-structured contextual information about an\nentity (such as a textual description, or visual depictions) the task is to\nextract the entity's underlying attributes. In this paper, we provide a dataset\ncontaining mixed-media data for over 2 million product items along with 7\nmillion attribute-value pairs describing the items which can be used to train\nattribute extractors in a weakly supervised manner. We provide a variety of\nbaselines which demonstrate the relative effectiveness of the individual modes\nof information towards solving the task, as well as study human performance.","url_abs":"http://arxiv.org/abs/1711.11118v1","url_pdf":"http://arxiv.org/pdf/1711.11118v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"multimodal-attribute-extraction","repo_url":"https://github.com/wavewangyue/mae","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"attribute","task_name":"Attribute"},{"task_slug":"attribute-extraction","task_name":"Attribute Extraction"},{"task_slug":"multimodal-attribute-value-extraction","task_name":"Multimodal Attribute Value Extraction"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1711.11118","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}