Papers › CheXGenBench: A Unified Benchmark For Fidelity, Privacy and Utility of Synthetic Chest...

CheXGenBench: A Unified Benchmark For Fidelity, Privacy and Utility of Synthetic Chest Radiographs

15 May 2025arXiv:2505.10496archive 2025-07-28

Raman Dutt, Pedro Sanchez, Yongchen Yao, Steven McDonagh, Sotirios A. Tsaftaris, Timothy Hospedales

We introduce CheXGenBench, a rigorous and multifaceted evaluation framework for synthetic chest radiograph generation that simultaneously assesses fidelity, privacy risks, and clinical utility across state-of-the-art text-to-image generative models. Despite rapid advancements in generative AI for real-world imagery, medical domain evaluations have been hindered by methodological inconsistencies, outdated architectural comparisons, and disconnected assessment criteria that rarely address the practical clinical value of synthetic samples. CheXGenBench overcomes these limitations through standardised data partitioning and a unified evaluation protocol comprising over 20 quantitative metrics that systematically analyse generation quality, potential privacy vulnerabilities, and downstream clinical applicability across 11 leading text-to-image architectures. Our results reveal critical inefficiencies in the existing evaluation protocols, particularly in assessing generative fidelity, leading to inconsistent and uninformative comparisons. Our framework establishes a standardised benchmark for the medical AI community, enabling objective and reproducible comparisons while facilitating seamless integration of both existing and future generative models. Additionally, we release a high-quality, synthetic dataset, SynthCheX-75K, comprising 75K radiographs generated by the top-performing model (Sana 0.6B) in our benchmark to support further research in this critical domain. Through CheXGenBench, we establish a new state-of-the-art and release our framework, models, and SynthCheX-75K dataset at https://raman1121.github.io/CheXGenBench/

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

Raman1121/CheXGenBench officialmentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Conditional Text-to-Image Synthesis

Datasets

Introduced by this paper, per the archive.

SynthCheX-75K

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Conditional Text-to-Image Synthesis MIMIC-CXR Sana FID (RadDino) 54.22 #1 of 13 Archive leaderboard report
Conditional Text-to-Image Synthesis MIMIC-CXR Pixart Sigma FID (RadDino) 60.15 #2 of 13 Archive leaderboard report
Conditional Text-to-Image Synthesis MIMIC-CXR RadEdit FID (RadDino) 69.69 #3 of 13 Archive leaderboard report
Conditional Text-to-Image Synthesis MIMIC-CXR LLM-CXR FID (RadDino) 71.24 #4 of 13 Archive leaderboard report
Conditional Text-to-Image Synthesis MIMIC-CXR SD V3.5 Medium (LoRA r128) FID (RadDino) 74.58 #5 of 13 Archive leaderboard report
Conditional Text-to-Image Synthesis MIMIC-CXR Lumina 2.0 (LoRA r128) FID (RadDino) 88.28 #6 of 13 Archive leaderboard report
Conditional Text-to-Image Synthesis MIMIC-CXR SD V3.5 Medium (LoRA r32) FID (RadDino) 93.10 #7 of 13 Archive leaderboard report
Conditional Text-to-Image Synthesis MIMIC-CXR Lumina 2.0 (LoRA r32) FID (RadDino) 101.19 #8 of 13 Archive leaderboard report
Conditional Text-to-Image Synthesis MIMIC-CXR SD V1-5 FID (RadDino) 118.93 #9 of 13 Archive leaderboard report
Conditional Text-to-Image Synthesis MIMIC-CXR Flux.1-Dev (LoRA r32) FID (RadDino) 122.40 #10 of 13 Archive leaderboard report
Conditional Text-to-Image Synthesis MIMIC-CXR SD V1-4 FID (RadDino) 125.18 #11 of 13 Archive leaderboard report
Conditional Text-to-Image Synthesis MIMIC-CXR SD V2-1 FID (RadDino) 186.53 #12 of 13 Archive leaderboard report
Conditional Text-to-Image Synthesis MIMIC-CXR SD V2 FID (RadDino) 194.72 #13 of 13 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections