{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/deep-multiple-instance-learning-for-zero-shot","title":"Deep Multiple Instance Learning for Zero-shot Image Tagging","arxiv_id":"1803.06051","date":"2018-03-16","proceeding":null,"authors":["Shafin Rahman","Salman Khan"],"abstract":"In-line with the success of deep learning on traditional recognition problem,\nseveral end-to-end deep models for zero-shot recognition have been proposed in\nthe literature. These models are successful to predict a single unseen label\ngiven an input image, but does not scale to cases where multiple unseen objects\nare present. In this paper, we model this problem within the framework of\nMultiple Instance Learning (MIL). To the best of our knowledge, we propose the\nfirst end-to-end trainable deep MIL framework for the multi-label zero-shot\ntagging problem. Due to its novel design, the proposed framework has several\ninteresting features: (1) Unlike previous deep MIL models, it does not use any\noff-line procedure (e.g., Selective Search or EdgeBoxes) for bag generation.\n(2) During test time, it can process any number of unseen labels given their\nsemantic embedding vectors. (3) Using only seen labels per image as weak\nannotation, it can produce a bounding box for each predicted labels. We\nexperiment with the NUS-WIDE dataset and achieve superior performance across\nconventional, zero-shot and generalized zero-shot tagging tasks.","url_abs":"http://arxiv.org/abs/1803.06051v1","url_pdf":"http://arxiv.org/pdf/1803.06051v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"deep-multiple-instance-learning-for-zero-shot","repo_url":"https://github.com/salman-h-khan/ZSD_Release","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"tf","reach":null}],"tasks":[{"task_slug":"multiple-instance-learning","task_name":"Multiple Instance Learning"},{"task_slug":"zero-shot-learning","task_name":"Zero-Shot Learning"}],"methods":[{"method_slug":"selective-search","method_name":"Selective Search"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1803.06051","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}