{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/hamlet-a-hierarchical-multimodal-attention-1","title":"HAMLET: A Hierarchical Multimodal Attention-based Human Activity Recognition Algorithm","arxiv_id":"2008.01148","date":"2020-08-03","proceeding":"International Conference on Intelligent Robots and Systems (IROS) 2020 6","authors":["Md Mofijul Islam","Tariq Iqbal"],"abstract":"To fluently collaborate with people, robots need the ability to recognize human activities accurately. Although modern robots are equipped with various sensors, robust human activity recognition (HAR) still remains a challenging task for robots due to difficulties related to multimodal data fusion. To address these challenges, in this work, we introduce a deep neural network-based multimodal HAR algorithm, HAMLET. HAMLET incorporates a hierarchical architecture, where the lower layer encodes spatio-temporal features from unimodal data by adopting a multi-head self-attention mechanism. We develop a novel multimodal attention mechanism for disentangling and fusing the salient unimodal features to compute the multimodal features in the upper layer. Finally, multimodal features are used in a fully connect neural-network to recognize human activities. We evaluated our algorithm by comparing its performance to several state-of-the-art activity recognition algorithms on three human activity datasets. The results suggest that HAMLET outperformed all other evaluated baselines across all datasets and metrics tested, with the highest top-1 accuracy of 95.12% and 97.45% on the UTD-MHAD [1] and the UT-Kinect [2] datasets respectively, and F1-score of 81.52% on the UCSD-MIT [3] dataset. We further visualize the unimodal and multimodal attention maps, which provide us with a tool to interpret the impact of attention mechanisms concerning HAR.","url_abs":"https://arxiv.org/abs/2008.01148v1","url_pdf":"https://arxiv.org/pdf/2008.01148v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"activity-recognition","task_name":"Activity Recognition"},{"task_slug":"human-activity-recognition","task_name":"Human Activity Recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/multimodal-activity-recognition-on-ucsd-mit","task":"Multimodal Activity Recognition","dataset":"UCSD-MIT Human Motion","model":"HAMLET","rank_in_archive_order":1,"of":1,"metrics":{"F1-score":"81.52"},"uses_additional_data":false},{"leaderboard":"/sota/multimodal-activity-recognition-on-ut-kinect","task":"Multimodal Activity Recognition","dataset":"UT-Kinect","model":"HAMLET","rank_in_archive_order":1,"of":1,"metrics":{"Accuracy (CS)":"97.56"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}