{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/hoi4abot-human-object-interaction","title":"HOI4ABOT: Human-Object Interaction Anticipation for Human Intention Reading Collaborative roBOTs","arxiv_id":"2309.16524","date":"2023-09-28","proceeding":null,"authors":["Esteve Valls Mascaro","Daniel Sliwowski","Dongheui Lee"],"abstract":"Robots are becoming increasingly integrated into our lives, assisting us in various tasks. To ensure effective collaboration between humans and robots, it is essential that they understand our intentions and anticipate our actions. In this paper, we propose a Human-Object Interaction (HOI) anticipation framework for collaborative robots. We propose an efficient and robust transformer-based model to detect and anticipate HOIs from videos. This enhanced anticipation empowers robots to proactively assist humans, resulting in more efficient and intuitive collaborations. Our model outperforms state-of-the-art results in HOI detection and anticipation in VidHOI dataset with an increase of 1.76% and 1.04% in mAP respectively while being 15.4 times faster. We showcase the effectiveness of our approach through experimental results in a real robot, demonstrating that the robot's ability to anticipate HOIs is key for better Human-Robot Interaction. More information can be found on our project webpage: https://evm7.github.io/HOI4ABOT_page/","url_abs":"https://arxiv.org/abs/2309.16524v2","url_pdf":"https://arxiv.org/pdf/2309.16524v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"human-object-interaction-anticipation","task_name":"Human-Object Interaction Anticipation"},{"task_slug":"human-object-interaction-detection","task_name":"Human-Object Interaction Detection"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/human-object-interaction-anticipation-on","task":"Human-Object Interaction Anticipation","dataset":"VidHOI","model":"HOI4ABOT","rank_in_archive_order":1,"of":3,"metrics":{"Person-wise Top5: t=1(mAP@0.5)":"37.77","Person-wise Top5: t=3(mAP@0.5)":"34.75","Person-wise Top5: t=5(mAP@0.5)":"34.07"},"uses_additional_data":false},{"leaderboard":"/sota/human-object-interaction-detection-on-vidhoi","task":"Human-Object Interaction Detection","dataset":"VidHOI","model":"HOI4ABOT","rank_in_archive_order":1,"of":3,"metrics":{"Detection: Full (mAP@0.5)":"11.12","Detection: Non-Rare (mAP@0.5)":"18.48","Detection: Rare (mAP@0.5)":"5.61","Oracle: Full (mAP@0.5)":"40.37","Oracle: Non-Rare (mAP@0.5)":"54.52","Oracle: Rare (mAP@0.5)":"29.5"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2309.16524","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}