{"url":"/dataset/simulacra-aesthetic-captions","name":"Simulacra Aesthetic Captions","full_name":null,"description_markdown":"**Simulacra Aesthetic Captions** is a dataset of over 238000 synthetic images generated with AI models such as CompVis latent GLIDE and Stable Diffusion from over forty thousand user submitted prompts. The images are rated on their aesthetic value from 1 to 10 by users to create caption, image, and rating triplets. In addition to this each user agreed to release all of their work with the bot: prompts, outputs, ratings, completely public domain under the CC0 1.0 Universal Public Domain Dedication. The result is a high quality royalty free dataset with over 176000 ratings that can be used for projects such as:\r\n\r\n- Filtering Datasets\r\n- Guiding Generative Models\r\n- Training A Prompt Generator\r\n- Extracting vitamin phrases (\"trending on artstation\", etc)\r\nAlignment Research\r\n\r\nDescription from: [https://github.com/JD-P/simulacra-aesthetic-captions](https://github.com/JD-P/simulacra-aesthetic-captions)","description_withheld":null,"homepage":"https://github.com/JD-P/simulacra-aesthetic-captions","introduced_date":"2022-07-03","introduced_date_note":null,"introduced_by":null,"license":{"name":"CC0 1.0 Universal Public Domain Dedication","url":"https://creativecommons.org/publicdomain/zero/1.0/"},"modalities":[{"name":"Images","url":"/datasets/modality/images"}],"tasks":[],"languages":[],"variants":["Simulacra Aesthetic Captions"],"data_loaders":[],"num_papers_in_archive":0,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-25T09:33:49+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}