{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/laion-400m-open-dataset-of-clip-filtered-400","title":"LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs","arxiv_id":"2111.02114","date":"2021-11-03","proceeding":null,"authors":["Christoph Schuhmann","Richard Vencu","Romain Beaumont","Robert Kaczmarczyk","Clayton Mullis","Aarush Katta","Theo Coombes","Jenia Jitsev","Aran Komatsuzaki"],"abstract":"Multi-modal language-vision models trained on hundreds of millions of image-text pairs (e.g. CLIP, DALL-E) gained a recent surge, showing remarkable capability to perform zero- or few-shot learning and transfer even in absence of per-sample labels on target image data. Despite this trend, to date there has been no publicly available datasets of sufficient scale for training such models from scratch. To address this issue, in a community effort we build and release for public LAION-400M, a dataset with CLIP-filtered 400 million image-text pairs, their CLIP embeddings and kNN indices that allow efficient similarity search.","url_abs":"https://arxiv.org/abs/2111.02114v1","url_pdf":"https://arxiv.org/pdf/2111.02114v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"laion-400m-open-dataset-of-clip-filtered-400","repo_url":"https://github.com/mlfoundations/open_clip","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"laion-400m-open-dataset-of-clip-filtered-400","repo_url":"https://github.com/compvis/latent-diffusion","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"laion-400m-open-dataset-of-clip-filtered-400","repo_url":"https://github.com/facebookresearch/metaclip","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"few-shot-learning","task_name":"Few-Shot Learning"}],"methods":[{"method_slug":"clip","method_name":"CLIP"}],"datasets_introduced":[{"slug":"laion-400m","name":"LAION-400M","full_name":""}],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=2111.02114","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2111.02114"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/mlfoundations/open_clip","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/facebookresearch/metaclip","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/compvis/latent-diffusion","reach":null}],"summary":{"unverified":1},"by_repo_kind":{"listed":{"samples":1,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":1,"samples":[{"code_sha256_prefix":"21f750ba403bf340","entry":"get_tokenizer","repo":"facebookresearch/metaclip","repo_kind":"listed","path":"src/mini_clip/factory.py","file_url":"https://github.com/facebookresearch/metaclip/blob/HEAD/src/mini_clip/factory.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"21f750ba403bf340"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}