{"url":"/dataset/twinsynths","name":"TwinSynths","full_name":null,"description_markdown":"The TwinSynths dataset is a novel benchmark designed to overcome common limitations found in earlier synthetic image datasets, such as low image quality, inadequate content preservation, and limited class diversity. TwinSynths generates pairs of images where each synthetic image is visually identical to its real counterpart, ensuring that the essential content remains intact while showcasing the unique architectural features of the generative models used. TwinSynths comprises two subsets:\r\n\r\n**TwinSynths-GAN** This subset uses a GAN generator architecture which trained from scratch on individual real images using a mean-squared error loss to ensure pixel-level fidelity. By fixing the latent vector input, the method produces synthetic images that closely mirror the original content. The GAN subset consists of 8,000 generated images spanning 80 classes selected from ImageNet.\r\n\r\n**TwinSynths-DM** For the diffusion model-based subset, DDIM inversion is used to maintain the content integrity of the original images. By applying a noise-adding forward process followed by a text-conditioned denoising procedure (using class name prompts), this generates synthetic images that are highly similar to their real counterparts. The same set of ImageNet classes is used as in the GAN subset, allowing for a consistent evaluation across different generative methods.","description_withheld":null,"homepage":"https://huggingface.co/datasets/koooooooook/TwinSynths","introduced_date":"2025-02-24","introduced_date_note":null,"introduced_by":{"paper":"/paper/sfld-reducing-the-content-bias-for-ai","title":"SFLD: Reducing the content bias for AI-generated Image Detection","first_author":"Seoyeon Gye","url":null},"license":null,"modalities":[{"name":"Images","url":"/datasets/modality/images"}],"tasks":[{"name":"Image Classification","url":"/task/image-classification","datasets_with_task":"/datasets/task/image-classification"},{"name":"Binary Classification","url":"/task/binary-classification","datasets_with_task":"/datasets/task/binary-classification"},{"name":"DeepFake Detection","url":"/task/deepfake-detection","datasets_with_task":"/datasets/task/deepfake-detection"},{"name":"Fake Image Detection","url":"/task/fake-image-detection","datasets_with_task":"/datasets/task/fake-image-detection"}],"languages":[],"variants":["TwinSynths"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}