Papers › SDAFE: A Dual-filter Stable Diffusion Data Augmentation Method for Facial Expression...

SDAFE: A Dual-filter Stable Diffusion Data Augmentation Method for Facial Expression Recognition

6 Apr 2025ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) 2025 4archive 2025-07-28

Minghao Zhao, Yifei Chen, Jiahao Lyu, Shuangli Du, Zhiyong Lv, Lin Wang

Facial expressions are a powerful medium for conveying emotions. In facial expression recognition (FER) field, the difficulty of collecting specific expressions often leads to class imbalance in mainstream datasets, significantly reducing the classification accuracy of deep neural networks. To address these issues, we propose a stable-diffusion-based augmentation method for facial expression (SDAFE) that resolves class imbalance problems and enhances data generation quality through cross-modal label guidance. By leveraging the neutrality of neutral faces, we generate additional expressions to balance the dataset classes. We introduce a peak signal-to-noise ratio (PSNR) filter to ensure the high quality of the generated images and a cosine similarity cross-modal filter based on CLIP encoders to ensure that the content of the generated images accurately aligns with their labels. Furthermore, we introduce a novel model, FERNeXt, which demonstrates outstanding performance in FER tasks, surpassing the state-of-the-art accuracy on the FER2013 dataset and achieving strong results on the RAF-DB and NHFI datasets. Subsequently, the performance of several models across different datasets significantly improves through the use of SDAFE in our experiments.

PaperPDF

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Data AugmentationFacial Expression RecognitionFacial Expression Recognition (FER)

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Facial Expression Recognition (FER) FER2013 FERNeXt-SDAFE Accuracy 81.33 #2 of 17 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

CLIP

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections