{"url":"/dataset/susy-dataset","name":"SuSy Dataset","full_name":null,"description_markdown":"The SuSy Dataset combines authentic photographs and AI-generated images designed for training and evaluating synthetic image detection models. It contains over 25,000 images from six different sources, including real-world photographs from COCO and synthetic images created by state-of-the-art diffusion models such as DALL-E 3, Midjourney, and Stable Diffusion.\r\n\r\n## Authentic Images\r\n- [COCO](https://cocodataset.org/) (Common Objects in Context): A large-scale object detection, segmentation, and captioning dataset. It includes over 330,000 images, with 200,000 labeled using 80 object categories. For this dataset, we use a random subset of 5,435 images. **License:** Creative Commons Attribution 4.0 license\r\n\r\n## Synthetic Images\r\n- [dalle-3-images](https://huggingface.co/datasets/ehristoforu/dalle-3-images): Contains 3,310 unique images generated using DALL-E 3. The dataset does not include the prompts used to generate the images. **License:** MIT license\r\n- [diffusiondb](https://poloclub.github.io/diffusiondb/): A large-scale text-to-image prompt dataset containing 14 million images generated by Stable Diffusion 1.x series models (2022). We use a random subset of 5,435 images. License:** CC0 1.0 Universal license\r\n- [realisticSDXL](https://huggingface.co/datasets/DucHaiten/DucHaiten-realistic-SDXL): Contains images generated using the Stable Diffusion XL (SDXL) model released in July 2023. We use only the \"realistic\" category, which contains 5,435 images. **License:** CreativeML OpenRAIL-M license\r\n- [midjourney-tti](https://www.kaggle.com/datasets/succinctlyai/midjourney-texttoimage): Contains images generated using Midjourney V1 or V2 models (early 2022). The original dataset provided URLs, which were scraped to obtain the images. **License:** CC0 1.0 Universal license (for links only, images are property of users who generated them)\r\n- [midjourney-images](https://huggingface.co/datasets/ehristoforu/midjourney-images): Contains 4,308 unique images generated using Midjourney V5 and V6 models (2023).  **License:** MIT license","description_withheld":null,"homepage":"https://huggingface.co/datasets/HPAI-BSC/SuSy-Dataset","introduced_date":"2024-09-21","introduced_date_note":null,"introduced_by":{"paper":"/paper/present-and-future-generalization-of","title":"Present and Future Generalization of Synthetic Image Detectors","first_author":"Pablo Bernabeu-Perez","url":null},"license":{"name":"Multiple licenses (see description)","url":null},"modalities":[{"name":"Images","url":"/datasets/modality/images"}],"tasks":[{"name":"Image Classification","url":"/task/image-classification","datasets_with_task":"/datasets/task/image-classification"},{"name":"Synthetic Image Detection","url":"/task/synthetic-image-detection","datasets_with_task":"/datasets/task/synthetic-image-detection"},{"name":"Synthetic Image Attribution","url":"/task/synthetic-image-attribution","datasets_with_task":"/datasets/task/synthetic-image-attribution"}],"languages":[],"variants":["SuSy Dataset"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-25T09:33:49+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}