{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-models-for-actions-and-person-object","title":"Learning Models for Actions and Person-Object Interactions with Transfer to Question Answering","arxiv_id":"1604.04808","date":"2016-04-16","proceeding":null,"authors":["Arun Mallya","Svetlana Lazebnik"],"abstract":"This paper proposes deep convolutional network models that utilize local and\nglobal context to make human activity label predictions in still images,\nachieving state-of-the-art performance on two recent datasets with hundreds of\nlabels each. We use multiple instance learning to handle the lack of\nsupervision on the level of individual person instances, and weighted loss to\nhandle unbalanced training data. Further, we show how specialized features\ntrained on these datasets can be used to improve accuracy on the Visual\nQuestion Answering (VQA) task, in the form of multiple choice fill-in-the-blank\nquestions (Visual Madlibs). Specifically, we tackle two types of questions on\nperson activity and person-object relationship and show improvements over\ngeneric features trained on the ImageNet classification task.","url_abs":"http://arxiv.org/abs/1604.04808v2","url_pdf":"http://arxiv.org/pdf/1604.04808v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"human-object-interaction-detection","task_name":"Human-Object Interaction Detection"},{"task_slug":"multiple-instance-learning","task_name":"Multiple Instance Learning"},{"task_slug":"multiple-choice","task_name":"Multiple-choice"},{"task_slug":"question-answering","task_name":"Question Answering"},{"task_slug":"visual-question-answering-1","task_name":"Visual Question Answering"},{"task_slug":"visual-question-answering","task_name":"Visual Question Answering (VQA)"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/human-object-interaction-detection-on-hico-1","task":"Human-Object Interaction Detection","dataset":"HICO","model":"Mallya & Lazebnik","rank_in_archive_order":6,"of":8,"metrics":{"mAP":"36.1"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1604.04808","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}