{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/on-the-power-of-data-augmentation-for-head","title":"On the power of data augmentation for head pose estimation","arxiv_id":"2407.05357","date":"2024-07-07","proceeding":null,"authors":["Michael Welter"],"abstract":"Deep learning has been impressively successful in the last decade in predicting human head poses from monocular images. However, for in-the-wild inputs the research community relies predominantly on a single training set, 300W-LP, of semisynthetic nature without many alternatives. This paper focuses on gradual extension and improvement of the data to explore the performance achievable with augmentation and synthesis strategies further. Modeling-wise a novel multitask head/loss design which includes uncertainty estimation is proposed. Overall, the thus obtained models are small, efficient, suitable for full 6 DoF pose estimation, and exhibit very competitive accuracy.","url_abs":"https://arxiv.org/abs/2407.05357v3","url_pdf":"https://arxiv.org/pdf/2407.05357v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"on-the-power-of-data-augmentation-for-head","repo_url":"https://github.com/opentrack/neuralnet-tracker-traincode","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"data-augmentation","task_name":"Data Augmentation"},{"task_slug":"face-alignment","task_name":"Face Alignment"},{"task_slug":"head-pose-estimation","task_name":"Head Pose Estimation"},{"task_slug":"pose-estimation","task_name":"Pose Estimation"}],"methods":[{"method_slug":"set","method_name":"SET"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/face-alignment-on-aflw2000-3d","task":"Face Alignment","dataset":"AFLW2000-3D","model":"OpNet","rank_in_archive_order":8,"of":14,"metrics":{"Balanced NME (2D Sparse Alignment)":"3.55%"},"uses_additional_data":true},{"leaderboard":"/sota/head-pose-estimation-on-aflw2000","task":"Head Pose Estimation","dataset":"AFLW2000","model":"OpNet","rank_in_archive_order":1,"of":25,"metrics":{"Geodesic Error (GE)":"5.23","MAE":"3.15"},"uses_additional_data":true},{"leaderboard":"/sota/head-pose-estimation-on-biwi","task":"Head Pose Estimation","dataset":"BIWI","model":"OpNet","rank_in_archive_order":8,"of":29,"metrics":{"Geodesic Error (GE)":"7.01","Geodesic Error - aligned (GE)":"4.72","MAE (trained with other data)":"3.57","MAE-aligned (trained with other data)":"2.65"},"uses_additional_data":true}],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}