{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/weakly-supervised-patchnets-describing-and","title":"Weakly Supervised PatchNets: Describing and Aggregating Local Patches for Scene Recognition","arxiv_id":"1609.00153","date":"2016-09-01","proceeding":null,"authors":["Zhe Wang","Li-Min Wang","Yali Wang","Bo-Wen Zhang","Yu Qiao"],"abstract":"Traditional feature encoding scheme (e.g., Fisher vector) with local\ndescriptors (e.g., SIFT) and recent convolutional neural networks (CNNs) are\ntwo classes of successful methods for image recognition. In this paper, we\npropose a hybrid representation, which leverages the discriminative capacity of\nCNNs and the simplicity of descriptor encoding schema for image recognition,\nwith a focus on scene recognition. To this end, we make three main\ncontributions from the following aspects. First, we propose a patch-level and\nend-to-end architecture to model the appearance of local patches, called {\\em\nPatchNet}. PatchNet is essentially a customized network trained in a weakly\nsupervised manner, which uses the image-level supervision to guide the\npatch-level feature extraction. Second, we present a hybrid visual\nrepresentation, called {\\em VSAD}, by utilizing the robust feature\nrepresentations of PatchNet to describe local patches and exploiting the\nsemantic probabilities of PatchNet to aggregate these local patches into a\nglobal representation. Third, based on the proposed VSAD representation, we\npropose a new state-of-the-art scene recognition approach, which achieves an\nexcellent performance on two standard benchmarks: MIT Indoor67 (86.2\\%) and\nSUN397 (73.0\\%).","url_abs":"http://arxiv.org/abs/1609.00153v2","url_pdf":"http://arxiv.org/pdf/1609.00153v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"weakly-supervised-patchnets-describing-and","repo_url":"https://github.com/wangzheallen/vsad","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"scene-recognition","task_name":"Scene Recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}