Papers › zGAN: An Outlier-focused Generative Adversarial Network For Realistic Synthetic Data Generation

zGAN: An Outlier-focused Generative Adversarial Network For Realistic Synthetic Data Generation

28 Oct 2024arXiv:2410.20808archive 2025-07-28

Azizjon Azimi, Bonu Boboeva, Ilyas Varshavskiy, Shuhrat Khalilbekov, Akhlitdin Nizamitdinov, Najima Noyoftova, Sergey Shulgin

The phenomenon of "black swans" has posed a fundamental challenge to performance of classical machine learning models. The perceived rise in frequency of outlier conditions, especially in post-pandemic environment, has necessitated exploration of synthetic data as a complement to real data in model training. This article provides a general overview and experimental investigation of the zGAN model architecture developed for the purpose of generating synthetic tabular data with outlier characteristics. The model is put to test in binary classification environments and shows promising results on realistic synthetic data generation, as well as uplift capabilities vis-\`a-vis model performance. A distinctive feature of zGAN is its enhanced correlation capability between features in the generated data, replicating correlations of features in real training data. Furthermore, crucial is the ability of zGAN to generate outliers based on covariance of real data or synthetically generated covariances. This approach to outlier generation enables modeling of complex economic events and augmentation of outliers for tasks such as training predictive models and detecting, processing or removing outliers. Experiments and comparative analyses as part of this study were conducted on both private (credit risk in financial services) and public datasets.

PaperPDF

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Binary ClassificationSynthetic Data EvaluationSynthetic Data GenerationSynthetic Outliers Evaluation

1 archive task tag without a task page not shown.

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Synthetic Data Evaluation Titanic zGAN AUC 0.8163 #1 of 7 Archive leaderboard report
Synthetic Data Evaluation Titanic CopulaGAN AUC 0.8076 #2 of 7 Archive leaderboard report
Synthetic Data Evaluation Titanic CTGAN AUC 0.7923 #3 of 7 Archive leaderboard report
Synthetic Data Evaluation Titanic TVAE AUC 0.7874 #4 of 7 Archive leaderboard report
Synthetic Data Evaluation Titanic SynthPop AUC 0.7861 #5 of 7 Archive leaderboard report
Synthetic Data Evaluation Titanic Gaussian Copula AUC 0.7846 #6 of 7 Archive leaderboard report
Synthetic Data Evaluation Titanic PrivBayes AUC 0.534 #7 of 7 Archive leaderboard report
Synthetic Outliers Evaluation A9 (3% outliers) zGAN AUC 0.7116 #1 of 1 Archive leaderboard report
Synthetic Outliers Evaluation A9 (5% outliers) zGAN AUC 0.7147 #1 of 1 Archive leaderboard report
Synthetic Outliers Evaluation A9 (7.4% outliers) zGAN AUC 0.7122 #1 of 1 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Introduced by this paper: Outlier Generation

Outlier Generation

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections