Papers › Analyzing Fairness in Deepfake Detection With Massively Annotated Databases

Analyzing Fairness in Deepfake Detection With Massively Annotated Databases

11 Aug 2022arXiv:2208.05845archive 2025-07-28

Ying Xu, Philipp Terhörst, Kiran Raja, Marius Pedersen

In recent years, image and video manipulations with Deepfake have become a severe concern for security and society. Many detection models and datasets have been proposed to detect Deepfake data reliably. However, there is an increased concern that these models and training databases might be biased and, thus, cause Deepfake detectors to fail. In this work, we investigate factors causing biased detection in public Deepfake datasets by (a) creating large-scale demographic and non-demographic attribute annotations with 47 different attributes for five popular Deepfake datasets and (b) comprehensively analysing attributes resulting in AI-bias of three state-of-the-art Deepfake detection backbone models on these datasets. The analysis shows how various attributes influence a large variety of distinctive attributes (from over 65M labels) on the detection performance which includes demographic (age, gender, ethnicity) and non-demographic (hair, skin, accessories, etc.) attributes. The results examined datasets show limited diversity and, more importantly, show that the utilised Deepfake detection backbone models are strongly affected by investigated attributes making them not fair across attributes. The Deepfake detection backbone methods trained on such imbalanced/biased datasets result in incorrect detection results leading to generalisability, fairness, and security issues. Our findings and annotated datasets will guide future research to evaluate and mitigate bias in Deepfake detection techniques. The annotated datasets and the corresponding code are publicly available.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

pterhoer/DeepFakeAnnotations officialmentioned in papermentioned on GitHub report
xuyingzhongguo/deepfakeannotations officialmentioned in papermentioned on GitHub report
leaffeall/DFS-GDD mentioned on GitHubpytorch report
purdue-m2/auc_fairness_with_noisy_groups mentioned on GitHubpytorchMIT report
purdue-m2/fairness-generalization mentioned on GitHubpytorchMIT report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

AttributeDecision MakingDeepFake DetectionFace SwappingFairness

Datasets

Introduced by this paper, per the archive.

DMAD

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections