Papers › Exposing flaws of generative model evaluation metrics and their unfair treatment of...
Exposing flaws of generative model evaluation metrics and their unfair treatment of diffusion models
George Stein, Jesse C. Cresswell, Rasa Hosseinzadeh, Yi Sui, Brendan Leigh Ross, Valentin Villecroze, Zhaoyan Liu, Anthony L. Caterini, J. Eric T. Taylor, Gabriel Loaiza-Ganem
We systematically study a wide variety of generative models spanning semantically-diverse image datasets to understand and improve the feature extractors and metrics used to evaluate them. Using best practices in psychophysics, we measure human perception of image realism for generated samples by conducting the largest experiment evaluating generative models to date, and find that no existing metric strongly correlates with human evaluations. Comparing to 17 modern metrics for evaluating the overall performance, fidelity, diversity, rarity, and memorization of generative models, we find that the state-of-the-art perceptual realism of diffusion models as judged by humans is not reflected in commonly reported metrics such as FID. This discrepancy is not explained by diversity in generated samples, though one cause is over-reliance on Inception-V3. We address these flaws through a study of alternative self-supervised feature extractors, find that the semantic information encoded by individual networks strongly depends on their training procedure, and show that DINOv2-ViT-L/14 allows for much richer evaluation of generative models. Next, we investigate data memorization, and find that generative models do memorize training examples on simple, smaller datasets like CIFAR10, but not necessarily on more complex datasets like ImageNet. However, our experiments show that current metrics do not properly detect memorization: none in the literature is able to separate memorization from other phenomena such as underfitting or mode shrinkage. To facilitate further development of generative models and their evaluation we release all generated image datasets, human evaluation data, and a modular library to compute 17 common metrics for 9 different encoders at https://github.com/layer6ai-labs/dgm-eval.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
For agents, Syntology's MCP tool lists every function and class Syntology harvested from this paper and whether it ran (how to connect): get_harvested_code_for_paper(arxiv_id="2306.04675")
Code
Syntology Ran 30 of 63 code samples harvested from 8 repositories linked to this paper; 33 have no recorded run. Of those that ran: 6 ran · honoured contract; 1 ran · violated contract; 14 ran · our draft was wrong; 5 ran · fixture could not drive it; 4 ran with no contract checked.
By repository: official repository: 3 samples from 1 repository, 2 ran; found in paper text by Syntology: 59 samples from 7 repositories, 27 ran; 1 identical to code first harvested elsewhere. The run record, sample by sample. “Ran” means executed on a synthesized input, not that the code is correct or reproduces the paper.
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
63 samples harvested; 30 ran; 6 honoured the contract we drafted; 33 have no recorded run. Read from Syntology's graph 2026-09-24; that is when this build read the record, not when the samples ran.
Licence: 1 of the 63 samples is pointer only, meaning Syntology does not serve that copy's text. This page shows no code text for any sample; each one links to its file in the repository.
Harvested from 8 repositories linked to this paper, official or community; each sample names its own and says which. Some samples are identical code Syntology first harvested from another repository; for those, this paper's copy is not located and its licence is not recorded. “Ran” means the sample executed on a synthesized input. It does not mean the output is correct, and nothing here reproduces the paper's results. “Honoured” and “violated” refer to a contract Syntology drafted from the code itself; “our draft was wrong” and “fixture could not drive it” are failures of Syntology's instrument, not of the code.
Each sample ends with its code_sha256, Syntology's identity for that exact code. An agent fetches the stored sample with Syntology's MCP tool get_code(code_sha256="…") (how to connect); click an identity to copy that call.
Repository labels, per sample. official repository: The archive marks this repository official for the paper. named in the paper: The archive records that the paper mentions this repository; it is not marked official. community (archive-listed): In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper. found in paper text by Syntology: Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted. community: Not in the archive's code links for this paper; a community repository Syntology harvested. Samples from a repository marked official are listed first. Licence labels name the repository's licence as recorded at harvest. “Pointer only” means Syntology does not serve that copy's text, for one of four reasons: no licence file was found; the licence was not identified; the licence is recorded as permissive but that copy's record is not marked cleared; or the licence is outside the permissive list Syntology serves text under (MIT, Apache-2.0, BSD and similar). Some licences outside that list permit redistribution, such as WTFPL, and GPL-3.0 under its conditions; they are simply not on the list. Hover a licence label for the reason. File links open the file on GitHub at the default branch, which may have changed since the harvest.
de1462c410a454f6 · report
0e74fdd296ca6e73 · report
2bbe398e90e76c2d · report
3e0a8dc623475d8e · report
9cd07419f7c926f8 · report
144d10b49baeb8a6 · report
ef1e7dffb5d433d4 · report
f87cd99314206f72 · report
0186912dcb8b92b9 · report
9b16691c6712e596 · report
85a4b5aac95f359d · report
70b3429a213d3584 · report
f6b944f50d3f15ae · report
a2b45afa097b55b6 · report
ffc38a27fd9393ba · report
7dfd4d71b4b83843 · report
456c172851946697 · report
6f65e378a4313f87 · report
f47dff4b7b27bba7 · report
dd1dc700e87f36cc · report
d01b68a634f71551 · report
236de8614a178382 · report
025e4c7cda618f25 · report
180016b9227381b1 · report
011230b2b9b8fb6f · report
5b0d8787e670fc63 · report
2dc23e3f121c6a38 · report
2168d015a2e457cb · report
fe99cfd294676edb · report
61bf152e6a42a184 · report
da4e9834e767ce2d · report
da365dda2e5af8e0 · report
1b2e0a7656cc76a7 · report
9d31d2c4cd16bb2d · report
9bd8ee32e810e900 · report
12b482fc759f87dd · report
0def50a351111624 · report
fc49fa553c005268 · report
258c1af08f41aa17 · report
053fc534bc6bb989 · report
77f4aa0649e7f404 · report
c1f9693596855557 · report
9865a02804b0266f · report
cd108d44d16c6e3c · report
8c4274c2f8994da8 · report
1983d19e3e1ab60d · report
6d3c813b7561493d · report
06f87811f5c2b96a · report
389c62669d4417ea · report
02870cd90e053ab0 · report
6a0f2095bdd516ab · report
e39a123f00eb58cd · report
95c51644bf3cb9ff · report
877afab065f859ba · report
d80134d1097de9c4 · report
d177a0713d0bea82 · report
38e696a22dc81a6f · report
0408240381821d2c · report
0f49a2b0642d1098 · report
43108d105300b7f5 · report
001e240d5a8fc4dc · report
5829218c5489c6b6 · report
8241c0562bc710fd · report
Tasks
Results from the paper archive 2025-07-28
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Methods
1 archive method tag without a method page not shown.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections