{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-visual-n-grams-from-web-data","title":"Learning Visual N-Grams from Web Data","arxiv_id":"1612.09161","date":"2016-12-29","proceeding":"ICCV 2017 10","authors":["Ang Li","Allan Jabri","Armand Joulin","Laurens van der Maaten"],"abstract":"Real-world image recognition systems need to recognize tens of thousands of\nclasses that constitute a plethora of visual concepts. The traditional approach\nof annotating thousands of images per class for training is infeasible in such\na scenario, prompting the use of webly supervised data. This paper explores the\ntraining of image-recognition systems on large numbers of images and associated\nuser comments. In particular, we develop visual n-gram models that can predict\narbitrary phrases that are relevant to the content of an image. Our visual\nn-gram models are feed-forward convolutional networks trained using new loss\nfunctions that are inspired by n-gram models commonly used in language\nmodeling. We demonstrate the merits of our models in phrase prediction,\nphrase-based image retrieval, relating images and captions, and zero-shot\ntransfer.","url_abs":"http://arxiv.org/abs/1612.09161v2","url_pdf":"http://arxiv.org/pdf/1612.09161v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"language-modeling","task_name":"Language Modeling"},{"task_slug":"language-modelling","task_name":"Language Modelling"},{"task_slug":"representation-learning","task_name":"Representation Learning"},{"task_slug":"retrieval","task_name":"Retrieval"},{"task_slug":"zero-shot-transfer-image-classification","task_name":"Zero-Shot Transfer Image Classification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/zero-shot-transfer-image-classification-on-2","task":"Zero-Shot Transfer Image Classification","dataset":"SUN","model":"Visual N-Grams","rank_in_archive_order":3,"of":3,"metrics":{"Accuracy":"23.0"},"uses_additional_data":false},{"leaderboard":"/sota/zero-shot-transfer-image-classification-on","task":"Zero-Shot Transfer Image Classification","dataset":"aYahoo","model":"Visual N-Grams","rank_in_archive_order":2,"of":2,"metrics":{"Accuracy":"72.4"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1612.09161","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}