Papers › FSA-Net: Learning Fine-Grained Structure Aggregation for Head Pose Estimation From a...
FSA-Net: Learning Fine-Grained Structure Aggregation for Head Pose Estimation From a Single Image
Tsun-Yi Yang, Yi-Ting Chen, Yen-Yu Lin, Yung-Yu Chuang
This paper proposes a method for head pose estimation from a single image. Previous methods often predict head poses through landmark or depth estimation and would require more computation than necessary. Our method is based on regression and feature aggregation. For having a compact model, we employ the soft stagewise regression scheme. Existing feature aggregation methods treat inputs as a bag of features and thus ignore their spatial relationship in a feature map. We propose to learn a fine-grained structure mapping for spatially grouping features before aggregation. The fine-grained structure provides part-based information and pooled values. By utilizing learnable and non-learnable importance over the spatial location, different model variants can be generated and form a complementary ensemble. Experiments show that our method outperforms the state-of-the-art methods including both the landmark-free ones and the ones based on landmark or depth estimation. With only a single RGB frame as input, our method even outperforms methods utilizing multi-modality information (RGB-D, RGB-Time) on estimating the yaw angle. Furthermore, the memory overhead of our model is 100 times smaller than those of previous methods.
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Head Pose Estimation | AFLW2000 | FSA-Net (Caps-Fusion) | Geodesic Error (GE) | 8.16 | #17 of 25 | Archive leaderboard | report |
| Head Pose Estimation | AFLW2000 | FSA-Net (Caps-Fusion) | MAE | 5.07 | #17 of 25 | Archive leaderboard | report |
| Head Pose Estimation | BIWI | FSA-Net (Caps-Fusion) | Geodesic Error (GE) | 7.64 | #13 of 29 | Archive leaderboard | report |
| Head Pose Estimation | BIWI | FSA-Net (Caps-Fusion) | Geodesic Error - aligned (GE) | 5.36 | #13 of 29 | Archive leaderboard | report |
| Head Pose Estimation | BIWI | FSA-Net (Caps-Fusion) | MAE (trained with other data) | 4.00 | #13 of 29 | Archive leaderboard | report |
| Head Pose Estimation | BIWI | FSA-Net (Caps-Fusion) | MAE-aligned (trained with other data) | 2.92 | #13 of 29 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections