{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/real-time-hand-gesture-recognition","title":"Real-Time Hand Gesture Recognition: Integrating Skeleton-Based Data Fusion and Multi-Stream CNN","arxiv_id":"2406.15003","date":"2024-06-21","proceeding":null,"authors":["Oluwaleke Yusuf","Maki Habib","Mohamed Moustafa"],"abstract":"Hand Gesture Recognition (HGR) enables intuitive human-computer interactions in various real-world contexts. However, existing frameworks often struggle to meet the real-time requirements essential for practical HGR applications. This study introduces a robust, skeleton-based framework for dynamic HGR that simplifies the recognition of dynamic hand gestures into a static image classification task, effectively reducing both hardware and computational demands. Our framework utilizes a data-level fusion technique to encode 3D skeleton data from dynamic gestures into static RGB spatiotemporal images. It incorporates a specialized end-to-end Ensemble Tuner (e2eET) Multi-Stream CNN architecture that optimizes the semantic connections between data representations while minimizing computational needs. Tested across five benchmark datasets (SHREC'17, DHG-14/28, FPHA, LMDHG, and CNR), the framework showed competitive performance with the state-of-the-art. Its capability to support real-time HGR applications was also demonstrated through deployment on standard consumer PC hardware, showcasing low latency and minimal resource usage in real-world settings. The successful deployment of this framework underscores its potential to enhance real-time applications in fields such as virtual/augmented reality, ambient intelligence, and assistive technologies, providing a scalable and efficient solution for dynamic gesture recognition.","url_abs":"https://arxiv.org/abs/2406.15003v2","url_pdf":"https://arxiv.org/pdf/2406.15003v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"real-time-hand-gesture-recognition","repo_url":"https://github.com/outsiders17711/e2eet-skeleton-based-hgr-using-data-level-fusion","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"gesture-recognition","task_name":"Gesture Recognition"},{"task_slug":"hand-gesture-recognition","task_name":"Hand Gesture Recognition"},{"task_slug":"hand-gesture-recognition-1","task_name":"Hand-Gesture Recognition"},{"task_slug":"image-classification","task_name":"Image Classification"},{"task_slug":"skeleton-based-action-recognition","task_name":"Skeleton Based Action Recognition"},{"task_slug":"image-classification","task_name":"image-classification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/hand-gesture-recognition-on-dhg-14","task":"Hand Gesture Recognition","dataset":"DHG-14","model":"e2eET","rank_in_archive_order":1,"of":13,"metrics":{"Accuracy":"95.83"},"uses_additional_data":false},{"leaderboard":"/sota/hand-gesture-recognition-on-dhg-28","task":"Hand Gesture Recognition","dataset":"DHG-28","model":"e2eET","rank_in_archive_order":2,"of":9,"metrics":{"Accuracy":"92.38"},"uses_additional_data":false},{"leaderboard":"/sota/hand-gesture-recognition-on-shrec-2017","task":"Hand Gesture Recognition","dataset":"SHREC 2017","model":"e2eET","rank_in_archive_order":1,"of":4,"metrics":{"14 Gestures Accuracy":"97.86","28 Gestures Accuracy":"95.36"},"uses_additional_data":false},{"leaderboard":"/sota/skeleton-based-action-recognition-on-first","task":"Skeleton Based Action Recognition","dataset":"First-Person Hand Action Benchmark","model":"e2eET","rank_in_archive_order":4,"of":4,"metrics":{"1:1 Accuracy":"91.83"},"uses_additional_data":false},{"leaderboard":"/sota/skeleton-based-action-recognition-on-sbu","task":"Skeleton Based Action Recognition","dataset":"SBU / SBU-Refine","model":"e2eET","rank_in_archive_order":8,"of":9,"metrics":{"Accuracy":"93.96"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}