{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/multimodal-autoregressive-pre-training-of","title":"Multimodal Autoregressive Pre-training of Large Vision Encoders","arxiv_id":"2411.14402","date":"2024-11-21","proceeding":"CVPR 2025 1","authors":["Enrico Fini","Mustafa Shukor","Xiujun Li","Philipp Dufter","Michal Klein","David Haldimann","Sai Aitharaju","Victor Guilherme Turrisi da Costa","Louis Béthune","Zhe Gan","Alexander T Toshev","Marcin Eichner","Moin Nabi","Yinfei Yang","Joshua M. Susskind","Alaaeldin El-Nouby"],"abstract":"We introduce a novel method for pre-training of large-scale vision encoders. Building on recent advancements in autoregressive pre-training of vision models, we extend this framework to a multimodal setting, i.e., images and text. In this paper, we present AIMV2, a family of generalist vision encoders characterized by a straightforward pre-training process, scalability, and remarkable performance across a range of downstream tasks. This is achieved by pairing the vision encoder with a multimodal decoder that autoregressively generates raw image patches and text tokens. Our encoders excel not only in multimodal evaluations but also in vision benchmarks such as localization, grounding, and classification. Notably, our AIMV2-3B encoder achieves 89.5% accuracy on ImageNet-1k with a frozen trunk. Furthermore, AIMV2 consistently outperforms state-of-the-art contrastive models (e.g., CLIP, SigLIP) in multimodal image understanding across diverse settings.","url_abs":"https://arxiv.org/abs/2411.14402v1","url_pdf":"https://arxiv.org/pdf/2411.14402v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"multimodal-autoregressive-pre-training-of","repo_url":"https://github.com/apple/ml-aim","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"jax","reach":{"status":"ok","spdx":"NOASSERTION"}}],"tasks":[{"task_slug":"decoder","task_name":"Decoder"},{"task_slug":"image-classification","task_name":"Image Classification"}],"methods":[{"method_slug":"clip","method_name":"CLIP"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/image-classification-on-imagenet","task":"Image Classification","dataset":"ImageNet","model":"AIMv2-3B (448 res)","rank_in_archive_order":18,"of":1060,"metrics":{"Top 1 Accuracy":"89.5%"},"uses_additional_data":false},{"leaderboard":"/sota/image-classification-on-imagenet","task":"Image Classification","dataset":"ImageNet","model":"AIMv2-3B","rank_in_archive_order":43,"of":1060,"metrics":{"Top 1 Accuracy":"88.5%"},"uses_additional_data":false},{"leaderboard":"/sota/image-classification-on-imagenet","task":"Image Classification","dataset":"ImageNet","model":"AIMv2-1B","rank_in_archive_order":60,"of":1060,"metrics":{"Number of params":"1200M","Top 1 Accuracy":"88.1%"},"uses_additional_data":false},{"leaderboard":"/sota/image-classification-on-imagenet","task":"Image Classification","dataset":"ImageNet","model":"AIMv2-H","rank_in_archive_order":87,"of":1060,"metrics":{"Number of params":"600M","Top 1 Accuracy":"87.5%"},"uses_additional_data":false},{"leaderboard":"/sota/image-classification-on-imagenet","task":"Image Classification","dataset":"ImageNet","model":"AIMv2-L","rank_in_archive_order":134,"of":1060,"metrics":{"Number of params":"300M","Top 1 Accuracy":"86.6%"},"uses_additional_data":false},{"leaderboard":"/sota/image-classification-on-imagenet","task":"Image Classification","dataset":"ImageNet","model":"AIMv2-2B","rank_in_archive_order":1059,"of":1060,"metrics":{"Number of params":"2700M"},"uses_additional_data":false},{"leaderboard":"/sota/image-classification-on-inaturalist","task":"Image Classification","dataset":"iNaturalist","model":"AIMv2-3B (448 res)","rank_in_archive_order":1,"of":19,"metrics":{"Top 1 Accuracy":"85.9"},"uses_additional_data":false},{"leaderboard":"/sota/image-classification-on-inaturalist","task":"Image Classification","dataset":"iNaturalist","model":"AIMv2-3B","rank_in_archive_order":5,"of":19,"metrics":{"Top 1 Accuracy":"81.5"},"uses_additional_data":false},{"leaderboard":"/sota/image-classification-on-inaturalist","task":"Image Classification","dataset":"iNaturalist","model":"AIMv2-1B","rank_in_archive_order":8,"of":19,"metrics":{"Top 1 Accuracy":"79.7"},"uses_additional_data":false},{"leaderboard":"/sota/image-classification-on-inaturalist","task":"Image Classification","dataset":"iNaturalist","model":"AIMv2-H","rank_in_archive_order":9,"of":19,"metrics":{"Top 1 Accuracy":"77.9"},"uses_additional_data":false},{"leaderboard":"/sota/image-classification-on-inaturalist","task":"Image Classification","dataset":"iNaturalist","model":"AIMv2-L","rank_in_archive_order":10,"of":19,"metrics":{"Top 1 Accuracy":"76"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2411.14402","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2411.14402"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/apple/ml-aim","reach":{"status":"ok","spdx":"NOASSERTION"}}],"summary":{"unverified":3},"by_repo_kind":{"official":{"samples":3,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":3,"samples":[{"code_sha256_prefix":"f2939a0126b2512e","entry":"accuracy","repo":"apple/ml-aim","repo_kind":"official","path":"aim-v1/aim/v1/utils.py","file_url":"https://github.com/apple/ml-aim/blob/HEAD/aim-v1/aim/v1/utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"f2939a0126b2512e"}},{"code_sha256_prefix":"21e85f13c1d7ea83","entry":"init_distributed_mode","repo":"apple/ml-aim","repo_kind":"official","path":"aim-v1/aim/v1/utils.py","file_url":"https://github.com/apple/ml-aim/blob/HEAD/aim-v1/aim/v1/utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"21e85f13c1d7ea83"}},{"code_sha256_prefix":"2c8b62c579ed1ba0","entry":"merge_state_dicts","repo":"apple/ml-aim","repo_kind":"official","path":"aim-v1/aim/v1/utils.py","file_url":"https://github.com/apple/ml-aim/blob/HEAD/aim-v1/aim/v1/utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"2c8b62c579ed1ba0"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}