{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/boosting-zero-shot-human-object-interaction","title":"Boosting Zero-Shot Human-Object Interaction Detection with Vision-Language Transfer","arxiv_id":null,"date":"2024-03-18","proceeding":"IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) 2024 3","authors":["Sandipan Sarma","Pradnesh Kalkar","Arijit Sur"],"abstract":"Human-Object Interaction (HOI) detection is a crucial task that involves localizing interactive human-object pairs and identifying the actions being performed. Most existing HOI detectors are supervised in nature and lack the ability of zero-shot discovery of unseen interactions. Recently, transformer-based methods have superseded the traditional CNN detectors by aggregating image-wide context but still suffer from the long-tail distribution problem in HOI. In this work, our primary focus is improving HOI detection in images, particularly in zero-shot scenarios. We use an end-to-end transformer-based object detector to localize human-object pairs and yield visual features of actions and objects. Moreover, we adopt the text encoder from a popular visual-language model called CLIP with a novel prompting mechanism to extract semantic information for unseen actions and objects. Finally, we learn a strong visual-semantic\r\nalignment and achieve state-of-the-art performance on the challenging HICO-DET dataset across five zero-shot settings, with up to 70.88% relative gains. Code is available at https://github.com/sandipan211/ZSHOI-VLT.","url_abs":"https://ieeexplore.ieee.org/document/10445910","url_pdf":"https://ieeexplore.ieee.org/document/10445910","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"boosting-zero-shot-human-object-interaction","repo_url":"https://github.com/sandipan211/ZSHOI-VLT","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"human-object-interaction-detection","task_name":"Human-Object Interaction Detection"},{"task_slug":"language-modeling","task_name":"Language Modeling"},{"task_slug":"language-modelling","task_name":"Language Modelling"},{"task_slug":"object","task_name":"Object"},{"task_slug":"zero-shot-human-object-interaction-detection","task_name":"Zero-Shot Human-Object Interaction Detection"}],"methods":[{"method_slug":"clip","method_name":"CLIP"},{"method_slug":"focus","method_name":"Focus"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}