{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/exploring-the-limits-of-weakly-supervised","title":"Exploring the Limits of Weakly Supervised Pretraining","arxiv_id":"1805.00932","date":"2018-05-02","proceeding":"ECCV 2018 9","authors":["Dhruv Mahajan","Ross Girshick","Vignesh Ramanathan","Kaiming He","Manohar Paluri","Yixuan Li","Ashwin Bharambe","Laurens van der Maaten"],"abstract":"State-of-the-art visual perception models for a wide range of tasks rely on\nsupervised pretraining. ImageNet classification is the de facto pretraining\ntask for these models. Yet, ImageNet is now nearly ten years old and is by\nmodern standards \"small\". Even so, relatively little is known about the\nbehavior of pretraining with datasets that are multiple orders of magnitude\nlarger. The reasons are obvious: such datasets are difficult to collect and\nannotate. In this paper, we present a unique study of transfer learning with\nlarge convolutional networks trained to predict hashtags on billions of social\nmedia images. Our experiments demonstrate that training for large-scale hashtag\nprediction leads to excellent results. We show improvements on several image\nclassification and object detection tasks, and report the highest ImageNet-1k\nsingle-crop, top-1 accuracy to date: 85.4% (97.6% top-5). We also perform\nextensive experiments that provide novel empirical data on the relationship\nbetween large-scale pretraining and transfer learning performance.","url_abs":"http://arxiv.org/abs/1805.00932v1","url_pdf":"http://arxiv.org/pdf/1805.00932v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"exploring-the-limits-of-weakly-supervised","repo_url":"https://github.com/eminorhan/resnext-wsl","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"exploring-the-limits-of-weakly-supervised","repo_url":"https://github.com/facebookresearch/ClassyVision","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"exploring-the-limits-of-weakly-supervised","repo_url":"https://github.com/facebookresearch/WSL-Images","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"NOASSERTION"}},{"paper_slug":"exploring-the-limits-of-weakly-supervised","repo_url":"https://github.com/PaddlePaddle/PaddleClas","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"paddle","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"image-classification","task_name":"Image Classification"},{"task_slug":"object-detection","task_name":"Object Detection"},{"task_slug":"transfer-learning","task_name":"Transfer Learning"},{"task_slug":"image-classification","task_name":"image-classification"},{"task_slug":"object-detection-1","task_name":"object-detection"}],"methods":[{"method_slug":"1x1-convolution","method_name":"1x1 Convolution"},{"method_slug":"average-pooling","method_name":"Average Pooling"},{"method_slug":"batch-normalization","method_name":"Batch Normalization"},{"method_slug":"convolution","method_name":"Convolution"},{"method_slug":"global-average-pooling","method_name":"Global Average Pooling"},{"method_slug":"grouped-convolution","method_name":"Grouped Convolution"},{"method_slug":"kaiming-initialization","method_name":"Kaiming Initialization"},{"method_slug":"randomhorizontalflip","method_name":"Random Horizontal Flip"},{"method_slug":"random-resized-crop","method_name":"Random Resized Crop"},{"method_slug":"relu","method_name":"ReLU"},{"method_slug":"resnext","method_name":"ResNeXt"},{"method_slug":"resnext-block","method_name":"ResNeXt Block"},{"method_slug":"residual-connection","method_name":"Residual Connection"},{"method_slug":"sgd-with-momentum","method_name":"SGD with Momentum"}],"datasets_introduced":[{"slug":"ig-3-5b-17k","name":"IG-3.5B-17k","full_name":"IG-3.5B-17k"}],"methods_introduced":[],"results":[{"leaderboard":"/sota/image-classification-on-imagenet","task":"Image Classification","dataset":"ImageNet","model":"ResNeXt-101 32x48d","rank_in_archive_order":236,"of":1060,"metrics":{"GFLOPs":"306","Number of params":"829M","Top 1 Accuracy":"85.4%"},"uses_additional_data":false},{"leaderboard":"/sota/image-classification-on-imagenet","task":"Image Classification","dataset":"ImageNet","model":"ResNeXt-101 32x32d","rank_in_archive_order":261,"of":1060,"metrics":{"GFLOPs":"174","Number of params":"466M","Top 1 Accuracy":"85.1%"},"uses_additional_data":false},{"leaderboard":"/sota/image-classification-on-imagenet","task":"Image Classification","dataset":"ImageNet","model":"ResNeXt-101 32×16d","rank_in_archive_order":345,"of":1060,"metrics":{"GFLOPs":"72","Number of params":"194M","Top 1 Accuracy":"84.2%"},"uses_additional_data":false},{"leaderboard":"/sota/image-classification-on-imagenet","task":"Image Classification","dataset":"ImageNet","model":"ResNeXt-101 32x8d","rank_in_archive_order":569,"of":1060,"metrics":{"Number of params":"88M","Top 1 Accuracy":"82.2%"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/1805.00932","atlas_url":"https://app.syntology.ai/?focus=1805.00932","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}