{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/the-unreasonable-effectiveness-of-noisy-data","title":"The Unreasonable Effectiveness of Noisy Data for Fine-Grained Recognition","arxiv_id":"1511.06789","date":"2015-11-20","proceeding":null,"authors":["Jonathan Krause","Benjamin Sapp","Andrew Howard","Howard Zhou","Alexander Toshev","Tom Duerig","James Philbin","Li Fei-Fei"],"abstract":"Current approaches for fine-grained recognition do the following: First,\nrecruit experts to annotate a dataset of images, optionally also collecting\nmore structured data in the form of part annotations and bounding boxes.\nSecond, train a model utilizing this data. Toward the goal of solving\nfine-grained recognition, we introduce an alternative approach, leveraging\nfree, noisy data from the web and simple, generic methods of recognition. This\napproach has benefits in both performance and scalability. We demonstrate its\nefficacy on four fine-grained datasets, greatly exceeding existing state of the\nart without the manual collection of even a single label, and furthermore show\nfirst results at scaling to more than 10,000 fine-grained categories.\nQuantitatively, we achieve top-1 accuracies of 92.3% on CUB-200-2011, 85.4% on\nBirdsnap, 93.4% on FGVC-Aircraft, and 80.8% on Stanford Dogs without using\ntheir annotated training sets. We compare our approach to an active learning\napproach for expanding fine-grained datasets.","url_abs":"http://arxiv.org/abs/1511.06789v3","url_pdf":"http://arxiv.org/pdf/1511.06789v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"the-unreasonable-effectiveness-of-noisy-data","repo_url":"https://github.com/google/goldfinch","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"NOASSERTION"}}],"tasks":[{"task_slug":"active-learning","task_name":"Active Learning"},{"task_slug":"fine-grained-image-classification","task_name":"Fine-Grained Image Classification"}],"methods":[],"datasets_introduced":[{"slug":"goldfinch","name":"Goldfinch","full_name":"GOogLe image-search Dataset"},{"slug":"l-bird","name":"L-Bird","full_name":"Large-Bird"}],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1511.06789","atlas_url":"https://app.syntology.ai/?focus=1511.06789","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}