Papers › Diffusion Models Beat GANs on Image Classification

Diffusion Models Beat GANs on Image Classification

17 Jul 2023arXiv:2307.08702archive 2025-07-28

Soumik Mukhopadhyay, Matthew Gwilliam, Vatsal Agarwal, Namitha Padmanabhan, Archana Swaminathan, Srinidhi Hegde, Tianyi Zhou, Abhinav Shrivastava

While many unsupervised learning models focus on one family of tasks, either generative or discriminative, we explore the possibility of a unified representation learner: a model which uses a single pre-training stage to address both families of tasks simultaneously. We identify diffusion models as a prime candidate. Diffusion models have risen to prominence as a state-of-the-art method for image generation, denoising, inpainting, super-resolution, manipulation, etc. Such models involve training a U-Net to iteratively predict and remove noise, and the resulting model can synthesize high fidelity, diverse, novel images. The U-Net architecture, as a convolution-based architecture, generates a diverse set of feature representations in the form of intermediate feature maps. We present our findings that these embeddings are useful beyond the noise prediction task, as they contain discriminative information and can also be leveraged for classification. We explore optimal methods for extracting and using these embeddings for classification tasks, demonstrating promising results on the ImageNet classification task. We find that with careful feature selection and pooling, diffusion models outperform comparable generative-discriminative methods such as BigBiGAN for classification tasks. We investigate diffusion models in the transfer learning regime, examining their performance on several fine-grained visual classification datasets. We compare these embeddings to those generated by competing architectures and pre-trainings for classification tasks.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

soumik-kanad/diffssl mentioned on GitHubpytorchMIT report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

ClassificationDenoisingFine-Grained Image ClassificationImage ClassificationImage GenerationSuper-ResolutionTransfer Learningfeature selectionimage-classification

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Methods

1x1 ConvolutionAdamAverage PoolingBatch NormalizationBigBiGANBigGANCReLUConcatenated Skip ConnectionConditional Batch NormalizationConvolutionDense ConnectionsDiffusionEarly StoppingFeature SelectionFeedforward NetworkFocusGAN Hinge LossGlobal Average PoolingLinear LayerMax PoolingNon-Local BlockNon-Local OperationOff-Diagonal Orthogonal RegularizationPointwise ConvolutionProjection DiscriminatorReLUResidual BlockResidual ConnectionRevNetReversible Residual BlockSAGANSoftmaxSpectral NormalizationTTURTruncation TrickU-Net

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections