{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/bilinear-cnns-for-fine-grained-visual","title":"Bilinear CNNs for Fine-grained Visual Recognition","arxiv_id":"1504.07889","date":"2015-04-29","proceeding":null,"authors":["Tsung-Yu Lin","Aruni RoyChowdhury","Subhransu Maji"],"abstract":"We present a simple and effective architecture for fine-grained visual\nrecognition called Bilinear Convolutional Neural Networks (B-CNNs). These\nnetworks represent an image as a pooled outer product of features derived from\ntwo CNNs and capture localized feature interactions in a translationally\ninvariant manner. B-CNNs belong to the class of orderless texture\nrepresentations but unlike prior work they can be trained in an end-to-end\nmanner. Our most accurate model obtains 84.1%, 79.4%, 86.9% and 91.3% per-image\naccuracy on the Caltech-UCSD birds [67], NABirds [64], FGVC aircraft [42], and\nStanford cars [33] dataset respectively and runs at 30 frames-per-second on a\nNVIDIA Titan X GPU. We then present a systematic analysis of these networks and\nshow that (1) the bilinear features are highly redundant and can be reduced by\nan order of magnitude in size without significant loss in accuracy, (2) are\nalso effective for other image classification tasks such as texture and scene\nrecognition, and (3) can be trained from scratch on the ImageNet dataset\noffering consistent improvements over the baseline architecture. Finally, we\npresent visualizations of these models on various datasets using top\nactivations of neural units and gradient-based inversion techniques. The source\ncode for the complete system is available at http://vis-www.cs.umass.edu/bcnn.","url_abs":"http://arxiv.org/abs/1504.07889v6","url_pdf":"http://arxiv.org/pdf/1504.07889v6.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"bilinear-cnns-for-fine-grained-visual","repo_url":"https://bitbucket.org/tsungyu/bcnn","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null},{"paper_slug":"bilinear-cnns-for-fine-grained-visual","repo_url":"https://github.com/haysacks/spot-the-coin","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}},{"paper_slug":"bilinear-cnns-for-fine-grained-visual","repo_url":"https://github.com/mangonihao/BiLinear_CNN","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"mindspore","reach":{"status":"ok"}},{"paper_slug":"bilinear-cnns-for-fine-grained-visual","repo_url":"https://github.com/tommarvoloriddle/Bilinear-CNN-Tensorflow2.4-implementation","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"fine-grained-image-classification","task_name":"Fine-Grained Image Classification"},{"task_slug":"fine-grained-visual-recognition","task_name":"Fine-Grained Visual Recognition"},{"task_slug":null,"task_name":"GPU"},{"task_slug":"image-classification","task_name":"Image Classification"},{"task_slug":"scene-recognition","task_name":"Scene Recognition"},{"task_slug":"image-classification","task_name":"image-classification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/fine-grained-image-classification-on-nabirds","task":"Fine-Grained Image Classification","dataset":"NABirds","model":"Bilinear-CNN","rank_in_archive_order":30,"of":30,"metrics":{"Accuracy":"79.4%"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1504.07889","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}