{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/multi-task-head-pose-estimation-in-the-wild-1","title":"Multi-task head pose estimation in-the-wild","arxiv_id":"2202.02299","date":"2020-12-22","proceeding":null,"authors":["Roberto Valle","José Miguel Buenaposada","Luis Baumela"],"abstract":"We present a deep learning-based multi-task approach for head pose estimation in images. We contribute with a network architecture and training strategy that harness the strong dependencies among face pose, alignment and visibility, to produce a top performing model for all three tasks. Our architecture is an encoder-decoder CNN with residual blocks and lateral skip connections. We show that the combination of head pose estimation and landmark-based face alignment significantly improve the performance of the former task. Further, the location of the pose task at the bottleneck layer, at the end of the encoder, and that of tasks depending on spatial information, such as visibility and alignment, in the final decoder layer, also contribute to increase the final performance. In the experiments conducted the proposed model outperforms the state-of-the-art in the face pose and visibility tasks. By including a final landmark regression step it also produces face alignment results on par with the state-of-the-art.","url_abs":"https://arxiv.org/abs/2202.02299v1","url_pdf":"https://arxiv.org/pdf/2202.02299v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"multi-task-head-pose-estimation-in-the-wild-1","repo_url":"https://github.com/bobetocalo/bobetocalo_pami20","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"tf","reach":null}],"tasks":[{"task_slug":"decoder","task_name":"Decoder"},{"task_slug":"face-alignment","task_name":"Face Alignment"},{"task_slug":"head-pose-estimation","task_name":"Head Pose Estimation"},{"task_slug":"pose-estimation","task_name":"Pose Estimation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/face-alignment-on-aflw2000","task":"Face Alignment","dataset":"AFLW2000","model":"MNN+ORB (Reannotated)","rank_in_archive_order":1,"of":5,"metrics":{"Error rate":"2.58"},"uses_additional_data":false},{"leaderboard":"/sota/face-alignment-on-aflw2000-3d","task":"Face Alignment","dataset":"AFLW2000-3D","model":"MNN+OR (reannotated)","rank_in_archive_order":1,"of":14,"metrics":{"Balanced NME (2D Sparse Alignment)":"2.58%"},"uses_additional_data":false},{"leaderboard":"/sota/face-alignment-on-cofw","task":"Face Alignment","dataset":"COFW","model":"MNN+OR (Inter-pupils Norm)","rank_in_archive_order":23,"of":28,"metrics":{"NME (inter-pupil)":"5.04%","Recall at 80% precision (Landmarks Visibility)":"72.12"},"uses_additional_data":false},{"leaderboard":"/sota/face-alignment-on-cofw","task":"Face Alignment","dataset":"COFW","model":"MNN (Inter-pupil Norm)","rank_in_archive_order":28,"of":28,"metrics":{"NME (inter-pupil)":"5.65%"},"uses_additional_data":false},{"leaderboard":"/sota/head-pose-estimation-on-aflw","task":"Head Pose Estimation","dataset":"AFLW","model":"MNN","rank_in_archive_order":1,"of":6,"metrics":{"MAE":"3.22"},"uses_additional_data":false},{"leaderboard":"/sota/head-pose-estimation-on-aflw2000","task":"Head Pose Estimation","dataset":"AFLW2000","model":"MNN","rank_in_archive_order":9,"of":25,"metrics":{"MAE":"3.83"},"uses_additional_data":false},{"leaderboard":"/sota/head-pose-estimation-on-biwi","task":"Head Pose Estimation","dataset":"BIWI","model":"MNN","rank_in_archive_order":9,"of":29,"metrics":{"MAE (trained with other data)":"3.66"},"uses_additional_data":false},{"leaderboard":"/sota/pose-estimation-on-300w-full","task":"Pose Estimation","dataset":"300W (Full)","model":"MNN","rank_in_archive_order":2,"of":3,"metrics":{"MAE mean (º)":"1.56"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=2202.02299","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}