Browse State-of-the-Art › Human-Object Interaction Detection
Human-Object Interaction Detection
173 papers with code · 6 benchmarks · 23 datasets archive 2025-07-28
Human-Object Interaction (HOI) detection is a task of identifying "a set of interactions" in an image, which involves the i) localization of the subject (i.e., humans) and target (i.e., objects) of interaction, and ii) the classification of the interaction labels.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
6 leaderboard tables shown for this task, 6 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| HICO-DET (55 rows) | Ours (PViC+) | Dynamic Scene Understanding from Vision-Language Representations | — | — | Compare |
| V-COCO (34 rows) | RLIPv2 | RLIPv2: Fast Scaling of Relational Language-Image Pre-training | code | Syntology ran 22 of 30 samples · 8 unverified | Compare |
| HICO (8 rows) | DEFR | The Overlooked Classifier in Human-Object Interaction Recognition | — | — | Compare |
| VidHOI (3 rows) | HOI4ABOT | HOI4ABOT: Human-Object Interaction Anticipation for Human... | — | — | Compare |
| Ambiguious-HOI (2 rows) | DJ-RN | Detailed 2D-3D Joint Representation for Human-Object Interaction | code | Syntology ran 0 of 12 samples · 12 unverified | Compare |
| MECCANO (1 row) | SlowFast + FasterRCNN | The MECCANO Dataset: Understanding Human-Object Interactions from... | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
23 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
2 subtasks in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 173 papers with code (449 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
22 Nov 2017 5 repositories listed Syntology ran 2 of 3 samples · 1 unverified · 3 pointer-only (licence)Temporal relational reasoning, the ability to link meaningful transformations of objects or entities over time, is a fundamental property of intelligent species.
-
24 Jul 2020 4 repositories listed Syntology ran 1 of 1 samples · 0 unverifiedThe integration of decomposition and composition enables VCL to share object and verb features among different HOI samples and images, and to generate new interaction samples and new types of HOI, and thus largely…
-
7 Jan 2020 4 repositories listedFew works have studied the disambiguating contribution of subsidiary relations made available via graph networks.
-
13 Apr 2019 4 repositories listedTo address these and promote the activity understanding, we build a large-scale Human Activity Knowledge Engine (HAKE) based on the human body part states.
-
30 Aug 2018 4 repositories listed Syntology ran 0 of 5 samples · 5 unverifiedOur core idea is that the appearance of a person or an object instance contains informative cues on which relevant parts of an image to attend to for facilitating interaction prediction.
-
18 Aug 2023 3 repositories listed Syntology ran 22 of 30 samples · 8 unverifiedIn this paper, we propose RLIPv2, a fast converging model that enables the scaling of relational pre-training to large-scale pseudo-labelled scene graph data.
-
5 Sep 2022 3 repositories listed Syntology ran 0 of 1 samples · 1 unverifiedThe task of Human-Object Interaction (HOI) detection targets fine-grained visual parsing of humans interacting with their environment, enabling a broad range of applications.
-
14 Feb 2022 3 repositories listed Syntology ran 0 of 1 samples · 1 unverifiedHuman activity understanding is of widespread interest in artificial intelligence and spans diverse applications like health care and behavior analysis.
-
26 Mar 2020 3 repositories listedOur quantitative and qualitative results show that (a) we can predict meaningful forces from videos whose effects lead to accurate imitation of the motions observed, (b) by jointly optimizing for contact point and force…
-
20 Nov 2018 3 repositories listed Syntology ran 0 of 6 samples · 6 unverifiedOn account of the generalization of interactiveness, interactiveness network is a transferable knowledge learner and can be cooperated with any HOI detection models to achieve desirable results.
-
14 Nov 2018 3 repositories listedWe show that for human-object interaction detection a relatively simple factorized model with appearance and layout encodings constructed from pre-trained object detectors outperforms more sophisticated approaches.
-
23 Oct 2023 2 repositories listed Syntology ran 4 of 8 samples · 4 unverifiedSpecifically, for predefined commonly used tag categories, RAM++ showcases 10.
-
26 Sep 2023 2 repositories listed Syntology ran 7 of 12 samples · 5 unverified · 12 pointer-only (licence)In contrast, we focus on inferring dense, 3D contact between the full body surface and objects in arbitrary images.
-
5 Jul 2023 2 repositories listedWe introduce a vision advisor decoder to fuse both the interaction region information and the VLM's vision knowledge and a Verb-HOI prediction bridge to promote interaction representation learning.
-
28 Aug 2022 2 repositories listed Syntology ran 3 of 8 samples · 5 unverifiedDue to the diversity of interactive affordance, the uniqueness of different individuals leads to diverse interactions, which makes it difficult to establish an explicit link between object parts and affordance labels.
-
27 Mar 2022 2 repositories listed Syntology ran 0 of 1 samples · 1 unverified · 1 pointer-only (licence)Therefore, the proposed method enables the learning on both known and unknown HOI concepts.
-
18 Mar 2022 2 repositories listedTo empower an agent with such ability, this paper proposes a task of affordance grounding from exocentric view, i.
-
19 Aug 2021 2 repositories listed Syntology ran 3 of 3 samples · 0 unverified · 3 pointer-only (licence)We evaluate this approach on our dataset, demonstrating that human-object relations can significantly reduce the ambiguity of articulated object reconstructions from challenging real-world videos.
-
7 Apr 2021 2 repositories listedThe proposed method can thus be used to 1) improve the performance of HOI detection, especially for the HOIs with unseen objects; and 2) infer the affordances of novel objects.
-
11 Dec 2020 2 repositories listedWe address the problem of detecting human-object interactions in images using graphical neural networks.
-
30 Oct 2020 2 repositories listedMeanwhile, isolated human and object can also be integrated into coherent HOI again.
-
14 Aug 2020 2 repositories listedWe consider the problem of Human-Object Interaction (HOI) Detection, which aims to locate and recognize HOI instances in the form of <human, action, object> in images.
-
7 Aug 2020 2 repositories listedTo address this issue, in this paper, we propose a novel Polysemy Deciphering Network (PD-Net) that decodes the visual polysemy of verbs for HOI detection in three distinct ways.
-
30 Jul 2020 2 repositories listedWe present a method that infers spatial arrangements and shapes of humans and objects in a globally consistent 3D scene, all from a single image in-the-wild captured in an uncontrolled environment.
-
2 Apr 2020 2 repositories listedIn light of this, we propose a new path: infer human part states first and then reason out the activities based on part-level semantics.
-
11 Mar 2020 2 repositories listed Syntology ran 0 of 8 samples · 8 unverifiedComprehensive visual understanding requires detection frameworks that can effectively learn and utilize object interactions while analyzing objects individually.
-
24 Apr 2017 2 repositories listedOur hypothesis is that the appearance of a person -- their pose, clothing, action -- is a powerful cue for localizing the objects they are interacting with.
-
5 May 2015 2 repositories listedIn this work, we exploit the simple observation that actions are accompanied by contextual cues to build a strong action recognition system.
-
12 Jul 2025 1 repository listedHuman-Object Interaction (HOI) detection is crucial for robot-human assistance, enabling context-aware support.
-
9 Jul 2025 1 repository listedThis framework includes an Attention Bias Guidance (ABG) component, which guides the VLM to produce fine-grained instance-level interaction features according to the attention bias provided by the HOI detector.
Syntology lines on 13 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections