{"url":"/dataset/artifact","name":"ArtiFact","full_name":"Artificial and Factual Image Dataset for Synthetic Image Detection","description_markdown":"The ArtiFact dataset is a large-scale image dataset that aims to include a diverse collection of real and synthetic images from multiple categories, including Human/Human Faces, Animal/Animal Faces, Places, Vehicles, Art, and many other real-life objects. The dataset comprises 8 sources that were carefully chosen to ensure diversity and includes images synthesized from 25 distinct methods, including 13 GANs, 7 Diffusion, and 5 other miscellaneous generators. The dataset contains 2,496,738 images, comprising 964,989 real images and 1,531,749 fake images.\r\n\r\nTo ensure diversity across different sources, the real images of the dataset are randomly sampled from source datasets containing numerous categories, whereas synthetic images are generated within the same categories as the real images. Captions and image masks from the COCO dataset are utilized to generate images for text2image and inpainting generators, while normally distributed noise with different random seeds is used for noise2image generators. The dataset is further processed to reflect real-world scenarios by applying random cropping, downscaling, and JPEG compression, in accordance with the [IEEE VIP Cup 2022 standards](https://grip-unina.github.io/vipcup2022/).\r\n\r\nThe ArtiFact dataset is intended to serve as a benchmark for evaluating the performance of synthetic image detectors under real-world conditions. It includes a broad spectrum of diversity in terms of generators used and syntheticity, providing a challenging dataset for image detection tasks.\r\n\r\n\r\n* Total number of images: 2,496,738\r\n* Number of real images: 964,989\r\n* Number of fake images: 1,531,749\r\n* Number of generators used for fake images: 25 (including 13 GANs, 7 Diffusion, and 5 miscellaneous generators)\r\n* Number of sources used for real images: 8\r\n* Categories included in the dataset: Human/Human Faces, Animal/Animal Faces, Places, Vehicles, Art, and other real-life objects\r\n* Image Resolution: 200 x 200","description_withheld":null,"homepage":"https://github.com/awsaf49/artifact","introduced_date":"2023-02-23","introduced_date_note":null,"introduced_by":{"paper":"/paper/artifact-a-large-scale-dataset-with","title":"ArtiFact: A Large-Scale Dataset with Artificial and Factual Images for Generalizable and Robust Synthetic Image Detection","first_author":"Md Awsafur Rahman","url":null},"license":null,"modalities":[{"name":"Images","url":"/datasets/modality/images"}],"tasks":[{"name":"Classification","url":"/task/classification-1","datasets_with_task":"/datasets/task/classification-1"},{"name":"Fake Image Detection","url":"/task/fake-image-detection","datasets_with_task":"/datasets/task/fake-image-detection"},{"name":"Fake Image Attribution","url":"/task/fake-image-attribution","datasets_with_task":"/datasets/task/fake-image-attribution"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["ArtiFact"],"data_loaders":[],"num_papers_in_archive":7,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}