Papers › Towards Accurate Facial Landmark Detection via Cascaded Transformers
Towards Accurate Facial Landmark Detection via Cascaded Transformers
Hui Li, Zidong Guo, Seon-Min Rhee, Seungju Han, Jae-Joon Han
Accurate facial landmarks are essential prerequisites for many tasks related to human faces. In this paper, an accurate facial landmark detector is proposed based on cascaded transformers. We formulate facial landmark detection as a coordinate regression task such that the model can be trained end-to-end. With self-attention in transformers, our model can inherently exploit the structured relationships between landmarks, which would benefit landmark detection under challenging conditions such as large pose and occlusion. During cascaded refinement, our model is able to extract the most relevant image features around the target landmark for coordinate prediction, based on deformable attention mechanism, thus bringing more accurate alignment. In addition, we propose a novel decoder that refines image features and landmark positions simultaneously. With few parameter increasing, the detection performance improves further. Our model achieves new state-of-the-art performance on several standard facial landmark detection benchmarks, and shows good generalization ability in cross-dataset evaluation.
In Syntology View this paper on Syntology: its repositories, every harvested function with whether it ran, its licence and the call to fetch it.
Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
No code repository is listed for this paper in the archive or in Syntology's graph.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Face Alignment | 300W | DTLD+ | NME_inter-ocular (%, Challenge) | 4.48 | #9 of 48 | Archive leaderboard | report |
| Face Alignment | 300W | DTLD+ | NME_inter-ocular (%, Common) | 2.6 | #9 of 48 | Archive leaderboard | report |
| Face Alignment | 300W | DTLD+ | NME_inter-ocular (%, Full) | 2.96 | #9 of 48 | Archive leaderboard | report |
| Face Alignment | 300W Split 2 | DTLD-s | AUC@7 (box) | 70.9 | #2 of 7 | Archive leaderboard | report |
| Face Alignment | 300W Split 2 | DTLD-s | NME (box) | 2.05 | #2 of 7 | Archive leaderboard | report |
| Face Alignment | AFLW-19 | DTLD+ | NME_diag (%, Full) | 1.37 | #6 of 23 | Archive leaderboard | report |
| Face Alignment | COFW | DTLD+ | NME (inter-ocular) | 3.02% | #3 of 28 | Archive leaderboard | report |
| Face Alignment | WFLW | DTLD+ | FR@10 (inter-ocular) | 2.68 | #5 of 36 | Archive leaderboard | report |
| Face Alignment | WFLW | DTLD+ | NME (inter-ocular) | 4.05 | #5 of 36 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections