{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/enhanced-spatio-temporal-image-encoding-for","title":"Enhanced Spatio- Temporal Image Encoding for Online Human Activity Recognition","arxiv_id":null,"date":"2023-12-15","proceeding":"International Conference on Machine Learning and Applications (ICMLA) 2023 12","authors":["Nassim Mokhtari","Vincent Fer","Alexis Nédélec","Marlene Gilles","Pierre De Loor"],"abstract":"Human Activity Recognition (HAR) based on sen-sors data can be seen as a time series classification problem where the challenge is to handle both spatial and temporal dependencies, while focusing on the most relevant data variations. It can be done using 3D skeleton data extracted from a RGB+D camera. In this work, we propose to improve the spatio-temporal image encoding of 3D skeletons captured from a Kinect sensor, by studying the concept of motion energy which focuses mainly on skeleton joints that are the most solicited for an action. This encoding allows us to achieve a better discrimination for the detection of online activities by focusing on the most significant parts of the actions. The article presents this new encoding and its application for HAR using a deep learning model trained on the encoded 3D skeleton data. For this purpose, we proposed to investigate the knowledge transferability of several pre-trained CNNs provided by Keras. The article shows a significant improvement of the accuracy of the learning according to the state of the art.","url_abs":"https://ieeexplore.ieee.org/abstract/document/10459847","url_pdf":"https://ieeexplore.ieee.org/abstract/document/10459847","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"enhanced-spatio-temporal-image-encoding-for","repo_url":"https://github.com/nassimmokhtari/Enhanced-Spatio-Temporal-Image-Encoding","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"activity-recognition","task_name":"Activity Recognition"},{"task_slug":"human-activity-recognition","task_name":"Human Activity Recognition"},{"task_slug":"time-series-1","task_name":"Time Series"},{"task_slug":"time-series-classification","task_name":"Time Series Classification"}],"methods":[{"method_slug":"vgg-16","method_name":"VGG-16"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/human-activity-recognition-on-oad-dataset","task":"Human Activity Recognition","dataset":"OAD dataset","model":"ESTIE + VGG16 (transfer-learning)","rank_in_archive_order":1,"of":3,"metrics":{"Accuracy":"95.22"},"uses_additional_data":false},{"leaderboard":"/sota/human-activity-recognition-on-oad-dataset","task":"Human Activity Recognition","dataset":"OAD dataset","model":"STIE + VGG16 (transfer-learning)","rank_in_archive_order":2,"of":3,"metrics":{"Accuracy":"94.77"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}