{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/convnext-v2-co-designing-and-scaling-convnets","title":"ConvNeXt V2: Co-designing and Scaling ConvNets with Masked Autoencoders","arxiv_id":"2301.00808","date":"2023-01-02","proceeding":"CVPR 2023 1","authors":["Sanghyun Woo","Shoubhik Debnath","Ronghang Hu","Xinlei Chen","Zhuang Liu","In So Kweon","Saining Xie"],"abstract":"Driven by improved architectures and better representation learning frameworks, the field of visual recognition has enjoyed rapid modernization and performance boost in the early 2020s. For example, modern ConvNets, represented by ConvNeXt, have demonstrated strong performance in various scenarios. While these models were originally designed for supervised learning with ImageNet labels, they can also potentially benefit from self-supervised learning techniques such as masked autoencoders (MAE). However, we found that simply combining these two approaches leads to subpar performance. In this paper, we propose a fully convolutional masked autoencoder framework and a new Global Response Normalization (GRN) layer that can be added to the ConvNeXt architecture to enhance inter-channel feature competition. This co-design of self-supervised learning techniques and architectural improvement results in a new model family called ConvNeXt V2, which significantly improves the performance of pure ConvNets on various recognition benchmarks, including ImageNet classification, COCO detection, and ADE20K segmentation. We also provide pre-trained ConvNeXt V2 models of various sizes, ranging from an efficient 3.7M-parameter Atto model with 76.7% top-1 accuracy on ImageNet, to a 650M Huge model that achieves a state-of-the-art 88.9% accuracy using only public training data.","url_abs":"https://arxiv.org/abs/2301.00808v1","url_pdf":"https://arxiv.org/pdf/2301.00808v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"convnext-v2-co-designing-and-scaling-convnets","repo_url":"https://github.com/facebookresearch/convnext-v2","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"convnext-v2-co-designing-and-scaling-convnets","repo_url":"https://github.com/rwightman/pytorch-image-models","is_official":0,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"convnext-v2-co-designing-and-scaling-convnets","repo_url":"https://github.com/IMvision12/keras-vision-models","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"convnext-v2-co-designing-and-scaling-convnets","repo_url":"https://github.com/Jacky-Android/convnext-v2-pytorch","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"convnext-v2-co-designing-and-scaling-convnets","repo_url":"https://github.com/Westlake-AI/openmixup","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"convnext-v2-co-designing-and-scaling-convnets","repo_url":"https://github.com/chenller/mmseg-extension","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"convnext-v2-co-designing-and-scaling-convnets","repo_url":"https://github.com/vishalned/MMEarth-train","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"convnext-v2-co-designing-and-scaling-convnets","repo_url":"https://github.com/zibbini/convnext-v2_tensorflow","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null},{"paper_slug":"convnext-v2-co-designing-and-scaling-convnets","repo_url":"https://github.com/2024-MindSpore-1/Code2/tree/main/model-1/convnextv2","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null},{"paper_slug":"convnext-v2-co-designing-and-scaling-convnets","repo_url":"https://github.com/Mind23-2/MindCode-22","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":{"status":"ok"}},{"paper_slug":"convnext-v2-co-designing-and-scaling-convnets","repo_url":"https://github.com/MindCode-4/code-3/tree/main/convnextv2","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null},{"paper_slug":"convnext-v2-co-designing-and-scaling-convnets","repo_url":"https://github.com/MindSpore-paper-code-3/code7/tree/main/ghostnetv2","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null},{"paper_slug":"convnext-v2-co-designing-and-scaling-convnets","repo_url":"https://github.com/leondgarse/keras_cv_attention_models/tree/main/keras_cv_attention_models/convnext","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"tf","reach":null},{"paper_slug":"convnext-v2-co-designing-and-scaling-convnets","repo_url":"https://github.com/lyqcom/convnext","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":{"status":"ok","spdx":"Apache-2.0"}},{"paper_slug":"convnext-v2-co-designing-and-scaling-convnets","repo_url":"https://github.com/mindspore-courses/External-Attention-MindSpore/blob/main/model/backbone/convnextv2.py","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null},{"paper_slug":"convnext-v2-co-designing-and-scaling-convnets","repo_url":"https://github.com/pwc-1/Paper-8/tree/main/convnext","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null},{"paper_slug":"convnext-v2-co-designing-and-scaling-convnets","repo_url":"https://gitlab.com/birder/birder","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"object-detection","task_name":"Object Detection"},{"task_slug":"representation-learning","task_name":"Representation Learning"},{"task_slug":"self-supervised-learning","task_name":"Self-Supervised Learning"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"}],"methods":[{"method_slug":"convnext","method_name":"ConvNeXt"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/semantic-segmentation-on-ade20k","task":"Semantic Segmentation","dataset":"ADE20K","model":"ConvNeXt V2-H (FCMAE)","rank_in_archive_order":51,"of":235,"metrics":{"Validation mIoU":"55"},"uses_additional_data":false},{"leaderboard":"/sota/semantic-segmentation-on-ade20k","task":"Semantic Segmentation","dataset":"ADE20K","model":"Swin V2-H","rank_in_archive_order":66,"of":235,"metrics":{"Validation mIoU":"54.2"},"uses_additional_data":false},{"leaderboard":"/sota/semantic-segmentation-on-ade20k","task":"Semantic Segmentation","dataset":"ADE20K","model":"ConvNeXt V2-L","rank_in_archive_order":76,"of":235,"metrics":{"Validation mIoU":"53.7"},"uses_additional_data":false},{"leaderboard":"/sota/semantic-segmentation-on-ade20k","task":"Semantic Segmentation","dataset":"ADE20K","model":"Swin-L","rank_in_archive_order":81,"of":235,"metrics":{"Validation mIoU":"53.5"},"uses_additional_data":false},{"leaderboard":"/sota/semantic-segmentation-on-ade20k","task":"Semantic Segmentation","dataset":"ADE20K","model":"Swin-B","rank_in_archive_order":87,"of":235,"metrics":{"Validation mIoU":"52.8"},"uses_additional_data":false},{"leaderboard":"/sota/semantic-segmentation-on-ade20k","task":"Semantic Segmentation","dataset":"ADE20K","model":"ConvNeXt V2-B","rank_in_archive_order":89,"of":235,"metrics":{"Validation mIoU":"52.1"},"uses_additional_data":false},{"leaderboard":"/sota/semantic-segmentation-on-ade20k","task":"Semantic Segmentation","dataset":"ADE20K","model":"ConvNeXt V2-L (Supervised)","rank_in_archive_order":94,"of":235,"metrics":{"Validation mIoU":"51.6"},"uses_additional_data":false},{"leaderboard":"/sota/semantic-segmentation-on-ade20k","task":"Semantic Segmentation","dataset":"ADE20K","model":"ConvNeXt V1-L","rank_in_archive_order":112,"of":235,"metrics":{"Validation mIoU":"50.5"},"uses_additional_data":false},{"leaderboard":"/sota/semantic-segmentation-on-ade20k","task":"Semantic Segmentation","dataset":"ADE20K","model":"ConvNeXt V1-B","rank_in_archive_order":125,"of":235,"metrics":{"Validation mIoU":"49.9"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=2301.00808","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2301.00808"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/rwightman/pytorch-image-models","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/Westlake-AI/openmixup","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/Mind23-2/MindCode-22","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/MindCode-4/code-3/tree/main/convnextv2","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/leondgarse/keras_cv_attention_models/tree/main/keras_cv_attention_models/convnext","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/mindspore-courses/External-Attention-MindSpore/blob/main/model/backbone/convnextv2.py","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/IMvision12/keras-vision-models","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/MindSpore-paper-code-3/code7/tree/main/ghostnetv2","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://gitlab.com/birder/birder","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/pwc-1/Paper-8/tree/main/convnext","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/Jacky-Android/convnext-v2-pytorch","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/vishalned/MMEarth-train","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/chenller/mmseg-extension","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/lyqcom/convnext","reach":{"status":"ok","spdx":"Apache-2.0"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/zibbini/convnext-v2_tensorflow","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/2024-MindSpore-1/Code2/tree/main/model-1/convnextv2","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/facebookresearch/convnext-v2","reach":null}],"summary":{"ran":10,"unverified":4},"by_repo_kind":{"official":{"samples":2,"ran":2,"repositories":1},"listed":{"samples":12,"ran":8,"repositories":4}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":9,"samples":[{"code_sha256_prefix":"2c6904dfd0e3ee2a","entry":"Block","repo":"vishalned/MMEarth-train","repo_kind":"listed","path":"models/convnextv2.py","file_url":"https://github.com/vishalned/MMEarth-train/blob/HEAD/models/convnextv2.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"2c6904dfd0e3ee2a"}},{"code_sha256_prefix":"af25ab15eb29c44b","entry":"Block","repo":"zibbini/convnext-v2_tensorflow","repo_kind":"listed","path":"convnext_pt/convnextv2.py","file_url":"https://github.com/zibbini/convnext-v2_tensorflow/blob/HEAD/convnext_pt/convnextv2.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"af25ab15eb29c44b"}},{"code_sha256_prefix":"f3c65efbbcd4b2a2","entry":"Block","repo":"Jacky-Android/convnext-v2-pytorch","repo_kind":"listed","path":"model.py","file_url":"https://github.com/Jacky-Android/convnext-v2-pytorch/blob/HEAD/model.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"f3c65efbbcd4b2a2"}},{"code_sha256_prefix":"24e783957539d177","entry":"Block","repo":"facebookresearch/convnext-v2","repo_kind":"official","path":"models/convnextv2.py","file_url":"https://github.com/facebookresearch/convnext-v2/blob/HEAD/models/convnextv2.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"invariant","behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"24e783957539d177"}},{"code_sha256_prefix":"d8c0f250ae1c4733","entry":"ConvNeXtV2","repo":"vishalned/MMEarth-train","repo_kind":"listed","path":"models/convnextv2.py","file_url":"https://github.com/vishalned/MMEarth-train/blob/HEAD/models/convnextv2.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"d8c0f250ae1c4733"}},{"code_sha256_prefix":"0b701ad2d565c3a4","entry":"ConvNeXtV2","repo":"zibbini/convnext-v2_tensorflow","repo_kind":"listed","path":"convnext_pt/convnextv2.py","file_url":"https://github.com/zibbini/convnext-v2_tensorflow/blob/HEAD/convnext_pt/convnextv2.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"0b701ad2d565c3a4"}},{"code_sha256_prefix":"420fed3739fee70b","entry":"ConvNeXtV2","repo":"facebookresearch/convnext-v2","repo_kind":"official","path":"models/convnextv2.py","file_url":"https://github.com/facebookresearch/convnext-v2/blob/HEAD/models/convnextv2.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"invariant","behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"420fed3739fee70b"}},{"code_sha256_prefix":"4035f346a2e3d70e","entry":"GRN","repo":"vishalned/MMEarth-train","repo_kind":"listed","path":"models/convnextv2.py","file_url":"https://github.com/vishalned/MMEarth-train/blob/HEAD/models/convnextv2.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"invariant","behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"4035f346a2e3d70e"}},{"code_sha256_prefix":"2f8121a6bb6bf49c","entry":"LayerNorm","repo":"vishalned/MMEarth-train","repo_kind":"listed","path":"models/convnextv2.py","file_url":"https://github.com/vishalned/MMEarth-train/blob/HEAD/models/convnextv2.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"2f8121a6bb6bf49c"}},{"code_sha256_prefix":"649912d2ea60c644","entry":"LayerNorm","repo":"Jacky-Android/convnext-v2-pytorch","repo_kind":"listed","path":"model.py","file_url":"https://github.com/Jacky-Android/convnext-v2-pytorch/blob/HEAD/model.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"649912d2ea60c644"}},{"code_sha256_prefix":"4b87ad6572b3686e","entry":"ConvNeXtV2","repo":"Jacky-Android/convnext-v2-pytorch","repo_kind":"listed","path":"model.py","file_url":"https://github.com/Jacky-Android/convnext-v2-pytorch/blob/HEAD/model.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"4b87ad6572b3686e"}},{"code_sha256_prefix":"177a4abf40724f6e","entry":"arg_to_varname","repo":"lyqcom/convnext","repo_kind":"listed","path":"src/configs/parser.py","file_url":"https://github.com/lyqcom/convnext/blob/HEAD/src/configs/parser.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"177a4abf40724f6e"}},{"code_sha256_prefix":"ffc9d7cced6f71fb","entry":"argv_to_vars","repo":"lyqcom/convnext","repo_kind":"listed","path":"src/configs/parser.py","file_url":"https://github.com/lyqcom/convnext/blob/HEAD/src/configs/parser.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"ffc9d7cced6f71fb"}},{"code_sha256_prefix":"a99f170ebd682af5","entry":"trim_preceding_hyphens","repo":"lyqcom/convnext","repo_kind":"listed","path":"src/configs/parser.py","file_url":"https://github.com/lyqcom/convnext/blob/HEAD/src/configs/parser.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"a99f170ebd682af5"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}