Papers › Multi-task head pose estimation in-the-wild
Multi-task head pose estimation in-the-wild
Roberto Valle, José Miguel Buenaposada, Luis Baumela
We present a deep learning-based multi-task approach for head pose estimation in images. We contribute with a network architecture and training strategy that harness the strong dependencies among face pose, alignment and visibility, to produce a top performing model for all three tasks. Our architecture is an encoder-decoder CNN with residual blocks and lateral skip connections. We show that the combination of head pose estimation and landmark-based face alignment significantly improve the performance of the former task. Further, the location of the pose task at the bottleneck layer, at the end of the encoder, and that of tasks depending on spatial information, such as visibility and alignment, in the final decoder layer, also contribute to increase the final performance. In the experiments conducted the proposed model outperforms the state-of-the-art in the face pose and visibility tasks. By including a final landmark regression step it also produces face alignment results on par with the state-of-the-art.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Face Alignment | AFLW2000 | MNN+ORB (Reannotated) | Error rate | 2.58 | #1 of 5 | Archive leaderboard | report |
| Face Alignment | AFLW2000-3D | MNN+OR (reannotated) | Balanced NME (2D Sparse Alignment) | 2.58% | #1 of 14 | Archive leaderboard | report |
| Face Alignment | COFW | MNN+OR (Inter-pupils Norm) | NME (inter-pupil) | 5.04% | #23 of 28 | Archive leaderboard | report |
| Face Alignment | COFW | MNN+OR (Inter-pupils Norm) | Recall at 80% precision (Landmarks Visibility) | 72.12 | #23 of 28 | Archive leaderboard | report |
| Face Alignment | COFW | MNN (Inter-pupil Norm) | NME (inter-pupil) | 5.65% | #28 of 28 | Archive leaderboard | report |
| Head Pose Estimation | AFLW | MNN | MAE | 3.22 | #1 of 6 | Archive leaderboard | report |
| Head Pose Estimation | AFLW2000 | MNN | MAE | 3.83 | #9 of 25 | Archive leaderboard | report |
| Head Pose Estimation | BIWI | MNN | MAE (trained with other data) | 3.66 | #9 of 29 | Archive leaderboard | report |
| Pose Estimation | 300W (Full) | MNN | MAE mean (º) | 1.56 | #2 of 3 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections