Papers › Using Self-Supervised Auxiliary Tasks to Improve Fine-Grained Facial Representation
Using Self-Supervised Auxiliary Tasks to Improve Fine-Grained Facial Representation
Mahdi Pourmirzaei, Gholam Ali Montazer, Farzaneh Esmaili
In this paper, at first, the impact of ImageNet pre-training on fine-grained Facial Emotion Recognition (FER) is investigated which shows that when enough augmentations on images are applied, training from scratch provides better result than fine-tuning on ImageNet pre-training. Next, we propose a method to improve fine-grained and in-the-wild FER, called Hybrid Multi-Task Learning (HMTL). HMTL uses Self-Supervised Learning (SSL) as an auxiliary task during classical Supervised Learning (SL) in the form of Multi-Task Learning (MTL). Leveraging SSL during training can gain additional information from images for the primary fine-grained SL task. We investigate how proposed HMTL can be used in the FER domain by designing two customized version of common pre-text task techniques, puzzling and in-painting. We achieve state-of-the-art results on the AffectNet benchmark via two types of HMTL, without utilizing pre-training on additional data. Experimental results on the common SSL pre-training and proposed HMTL demonstrate the difference and superiority of our work. However, HMTL is not only limited to FER domain. Experiments on two types of fine-grained facial tasks, i.e., head pose estimation and gender recognition, reveals the potential of using HMTL to improve fine-grained facial representation.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
No code repository is listed for this paper in the archive or in Syntology's graph.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Facial Expression Recognition (FER) | AffectNet | SL + SSL in-panting-pl (B0) | Accuracy (8 emotion) | 61.72 | #17 of 50 | Archive leaderboard | report |
| Facial Expression Recognition (FER) | AffectNet | SL + SSL puzzling (B2) | Accuracy (8 emotion) | 61.32 | #20 of 50 | Archive leaderboard | report |
| Facial Expression Recognition (FER) | AffectNet | SL + SSL puzzling (B0) | Accuracy (8 emotion) | 61.09 | #21 of 50 | Archive leaderboard | report |
| Facial Expression Recognition (FER) | AffectNet | SL (B2) | Accuracy (8 emotion) | 60.35 | #24 of 50 | Archive leaderboard | report |
| Facial Expression Recognition (FER) | AffectNet | SL (B0) | Accuracy (8 emotion) | 60.34 | #25 of 50 | Archive leaderboard | report |
| Facial Expression Recognition (FER) | AffectNet | SL+ SSL in-painting-pl + 20% train (B0) | Accuracy (8 emotion) | 55.36 | #34 of 50 | Archive leaderboard | report |
| Facial Expression Recognition (FER) | AffectNet | SL+ SSL puzzling + 20% train (B0) | Accuracy (8 emotion) | 54.98 | #35 of 50 | Archive leaderboard | report |
| Facial Expression Recognition (FER) | AffectNet | SL + 20% train (B0) | Accuracy (8 emotion) | 52.46 | #37 of 50 | Archive leaderboard | report |
| Facial Expression Recognition (FER) | CK+ | Nonlinear eval on SL + SSL puzzling (B0) | Accuracy (7 emotion) | 98.23 | #6 of 7 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Methods
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections