Home › Datasets › task › Human-Object Interaction Detection

Human-Object Interaction Detection datasets

archive 2025-07-28

23 datasets carry the task tag "Human-Object Interaction Detection" (the task itself: Human-Object Interaction Detection), ordered by the archive's paper count. Page 1 of 1: 23 shown of 23. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Human-Object Interaction Detection datasets 1–23 of 23

V-COCO (Verbs in COCO)
Verbs in COCO (V-COCO) is a dataset that builds off COCO for human-object interaction detection.
159 papers · 1 benchmark
FineGym is an action recognition dataset build on top of gymnasium videos.
76 papers · 0 benchmarks
BEHAVE is a full body human-object interaction dataset with multi-view RGBD frames and corresponding 3D SMPL and object fits along with the annotated contacts between them.
53 papers · 3 benchmarks
RICH (Real scenes, Interaction, Contact and Humans)
Inferring human-scene contact (HSC) is the first step toward understanding how humans interact with their surroundings.
53 papers · 1 benchmark
HICO (Humans Interacting with Common Objects)
HICO is a benchmark for recognizing human-object interactions (HOI).
45 papers · 1 benchmark
A large-scale 4D egocentric dataset with rich annotations, to catalyze the research of category-level human-object interaction.
23 papers · 0 benchmarks
The MECCANO dataset is the first dataset of egocentric videos to study human-object interactions in industrial-like settings.
19 papers · 3 benchmarks
HAKE is built upon existing activity datasets and provides human body part level atomic action labels (Part States).
15 papers · 0 benchmarks
The Watch-n-Patch dataset was created with the focus on modeling human activities, comprising multiple actions in a completely unsupervised setting.
13 papers · 0 benchmarks
COUCH is a large human-chair interaction dataset with clean annotations.
8 papers · 0 benchmarks
VidHOI is a video-based human-object interaction detection benchmark.
7 papers · 2 benchmarks
CHAIRS is a large-scale motion-captured f-AHOI dataset, consisting of 17.3 hours of versatile interactions between 46 participants and 81 articulated and rigid sittable objects.
3 papers · 0 benchmarks
V-HICO is a dataset for human-object interaction in videos.
3 papers · 0 benchmarks
Ambiguous-HOI is a challenging dataset containing ambiguous human-object interaction images for HOI detection based on HICO-DET.
2 papers · 0 benchmarks
MPHOI-72 (Multi-person Human-object Interaction Dataset 72)
MPHOI-72 is a multi-person human-object interaction dataset that can be used for a wide variety of HOI/activity recognition and pose estimation/object tracking tasks.
2 papers · 0 benchmarks
DARai (Daily Activity Recordings for AI and ML applications)
Daily Activity Recordings for Artificial Intelligence (DARai, pronounced "Dahr-ree") is a multimodal, hierarchically annotated dataset constructed to understand human activities in real-world settings.
1 paper · 0 benchmarks
DIO (Discovering Interacted Objects)
Discovering Interacted Objects (DIO) is a benchmark containing 51 interactions and 1,000+ objects designed for Spatio-temporal Human-Object Interaction (ST-HOI) detection.
1 paper · 0 benchmarks
EgoISM-HOI is a new multimodal dataset composed of synthetic and real images of egocentric human-objects interactions in an industrial environment with rich annotations of hands and objects.
1 paper · 0 benchmarks
The Human-to-Human-or-Object Interaction Dataset (H2O) dataset is a dataset for Human-Object Interaction (HOI) detection.
1 paper · 0 benchmarks
The dataset is composed of 100 video sequences densely annotated with 60K bounding boxes, 17 sequence attributes, 13 action verb attributes and 29 target object attributes.
1 paper · 0 benchmarks
First of its kind paired win-fail action understanding dataset with samples from the following domains: “General Stunts,” “Internet Wins-Fails,” “Trick Shots,” & “Party Games.” The task is to identify successful and failed attempts at…
1 paper · 2 benchmarks
H²O Interaction (Human-to-Human-or-Object Interaction)
H²O is an image dataset annotated for Human-to-human-or-object interaction detection.
0 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.