{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/exploit-bounding-box-annotations-for-multi","title":"Exploit Bounding Box Annotations for Multi-label Object Recognition","arxiv_id":"1504.05843","date":"2015-04-22","proceeding":"CVPR 2016 6","authors":["Hao Yang","Joey Tianyi Zhou","Yu Zhang","Bin-Bin Gao","Jianxin Wu","Jianfei Cai"],"abstract":"Convolutional neural networks (CNNs) have shown great performance as general\nfeature representations for object recognition applications. However, for\nmulti-label images that contain multiple objects from different categories,\nscales and locations, global CNN features are not optimal. In this paper, we\nincorporate local information to enhance the feature discriminative power. In\nparticular, we first extract object proposals from each image. With each image\ntreated as a bag and object proposals extracted from it treated as instances,\nwe transform the multi-label recognition problem into a multi-class\nmulti-instance learning problem. Then, in addition to extracting the typical\nCNN feature representation from each proposal, we propose to make use of\nground-truth bounding box annotations (strong labels) to add another level of\nlocal information by using nearest-neighbor relationships of local regions to\nform a multi-view pipeline. The proposed multi-view multi-instance framework\nutilizes both weak and strong labels effectively, and more importantly it has\nthe generalization ability to even boost the performance of unseen categories\nby partial strong labels from other categories. Our framework is extensively\ncompared with state-of-the-art hand-crafted feature based methods and CNN based\nmethods on two multi-label benchmark datasets. The experimental results\nvalidate the discriminative power and the generalization ability of the\nproposed framework. With strong labels, our framework is able to achieve\nstate-of-the-art results in both datasets.","url_abs":"http://arxiv.org/abs/1504.05843v2","url_pdf":"http://arxiv.org/pdf/1504.05843v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"multi-label-classification","task_name":"Multi-Label Classification"},{"task_slug":"object","task_name":"Object"},{"task_slug":"object-recognition","task_name":"Object Recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/multi-label-classification-on-pascal-voc-2007","task":"Multi-Label Classification","dataset":"PASCAL VOC 2007","model":"FeV+LV  (pretrain from ImageNet)","rank_in_archive_order":17,"of":17,"metrics":{"mAP":"92.0"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1504.05843","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}