{"url":"/dataset/laion-5b","name":"LAION-5B","full_name":null,"description_markdown":"**LAION 5B** is a large-scale dataset for research purposes consisting of 5,85B CLIP-filtered image-text pairs. 2,3B contain English language, 2,2B samples from 100+ other languages and 1B samples have texts that do not allow a certain language assignment (e.g. names ). Additionally, we provide several nearest neighbor indices, an improved web interface for exploration & subset creation as well as detection scores for watermark and NSFW.","description_withheld":null,"homepage":"https://laion.ai/blog/laion-5b/","introduced_date":"2022-10-16","introduced_date_note":null,"introduced_by":{"paper":"/paper/laion-5b-an-open-large-scale-dataset-for-1","title":"LAION-5B: An open large-scale dataset for training next generation image-text models","first_author":"Christoph Schuhmann","url":null},"license":{"name":"Creative Common CC-BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"modalities":[{"name":"Images","url":"/datasets/modality/images"},{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[],"languages":[],"variants":["LAION-5B"],"data_loaders":[],"num_papers_in_archive":205,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}