{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/simple-yet-efficient-real-time-pose-based","title":"Simple yet efficient real-time pose-based action recognition","arxiv_id":"1904.09140","date":"2019-04-19","proceeding":null,"authors":["Dennis Ludl","Thomas Gulde","Cristóbal Curio"],"abstract":"Recognizing human actions is a core challenge for autonomous systems as they\ndirectly share the same space with humans. Systems must be able to recognize\nand assess human actions in real-time. In order to train corresponding\ndata-driven algorithms, a significant amount of annotated training data is\nrequired. We demonstrated a pipeline to detect humans, estimate their pose,\ntrack them over time and recognize their actions in real-time with standard\nmonocular camera sensors. For action recognition, we encode the human pose into\na new data format called Encoded Human Pose Image (EHPI) that can then be\nclassified using standard methods from the computer vision community. With this\nsimple procedure we achieve competitive state-of-the-art performance in\npose-based action detection and can ensure real-time performance. In addition,\nwe show a use case in the context of autonomous driving to demonstrate how such\na system can be trained to recognize human actions using simulation data.","url_abs":"http://arxiv.org/abs/1904.09140v1","url_pdf":"http://arxiv.org/pdf/1904.09140v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"simple-yet-efficient-real-time-pose-based","repo_url":"https://github.com/noboevbo/ehpi_action_recognition","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"action-detection","task_name":"Action Detection"},{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"autonomous-driving","task_name":"Autonomous Driving"},{"task_slug":"skeleton-based-action-recognition","task_name":"Skeleton Based Action Recognition"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/skeleton-based-action-recognition-on-j-hmdb","task":"Skeleton Based Action Recognition","dataset":"J-HMDB","model":"EHPI","rank_in_archive_order":13,"of":13,"metrics":{"Accuracy (RGB+pose)":"-","Accuracy (pose)":"65.5"},"uses_additional_data":false},{"leaderboard":"/sota/skeleton-based-action-recognition-on-jhmdb-2d","task":"Skeleton Based Action Recognition","dataset":"JHMDB (2D poses only)","model":"EHPI","rank_in_archive_order":5,"of":6,"metrics":{"Average accuracy of 3 splits":"65.5"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}