Papers › Cascaded Dual Vision Transformer for Accurate Facial Landmark Detection
Cascaded Dual Vision Transformer for Accurate Facial Landmark Detection
Ziqiang Dang, Jianfang Li, Lin Liu
Facial landmark detection is a fundamental problem in computer vision for many downstream applications. This paper introduces a new facial landmark detector based on vision transformers, which consists of two unique designs: Dual Vision Transformer (D-ViT) and Long Skip Connections (LSC). Based on the observation that the channel dimension of feature maps essentially represents the linear bases of the heatmap space, we propose learning the interconnections between these linear bases to model the inherent geometric relations among landmarks via Channel-split ViT. We integrate such channel-split ViT into the standard vision transformer (i.e., spatial-split ViT), forming our Dual Vision Transformer to constitute the prediction blocks. We also suggest using long skip connections to deliver low-level image features to all prediction blocks, thereby preventing useful information from being discarded by intermediate supervision. Extensive experiments are conducted to evaluate the performance of our proposal on the widely used benchmarks, i.e., WFLW, COFW, and 300W, demonstrating that our model outperforms the previous SOTAs across all three benchmarks.
In Syntology View this paper on Syntology: its repositories, every harvested function with whether it ran, its licence and the call to fetch it.
Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Facial Landmark Detection | 300W | D-ViT | NME | 2.85 | #1 of 15 | Archive leaderboard | report |
| Facial Landmark Detection | COFW | D-ViT | NME (inter-pupil) | 4.13 | #1 of 2 | Archive leaderboard | report |
| Facial Landmark Detection | WFLW | D-ViT | AUC@10 (inter-ocular) | 63.7 | #1 of 3 | Archive leaderboard | report |
| Facial Landmark Detection | WFLW | D-ViT | FR@10 (inter-ocular) | 1.76 | #1 of 3 | Archive leaderboard | report |
| Facial Landmark Detection | WFLW | D-ViT | NME | 3.75 | #1 of 3 | Archive leaderboard | report |
| Facial Landmark Detection | WFLW | D-ViT | NME (inter-ocular) | 3.75 | #1 of 3 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Methods
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections