Papers › Bongard-HOI: Benchmarking Few-Shot Visual Reasoning for Human-Object Interactions

Bongard-HOI: Benchmarking Few-Shot Visual Reasoning for Human-Object Interactions

27 May 2022CVPR 2022 1arXiv:2205.13803archive 2025-07-28

Huaizu Jiang, Xiaojian Ma, Weili Nie, Zhiding Yu, Yuke Zhu, Song-Chun Zhu, Anima Anandkumar

A significant gap remains between today's visual pattern recognition models and human-level visual cognition especially when it comes to few-shot learning and compositional reasoning of novel concepts. We introduce Bongard-HOI, a new visual reasoning benchmark that focuses on compositional learning of human-object interactions (HOIs) from natural images. It is inspired by two desirable characteristics from the classical Bongard problems (BPs): 1) few-shot concept learning, and 2) context-dependent reasoning. We carefully curate the few-shot instances with hard negatives, where positive and negative images only disagree on action labels, making mere recognition of object categories insufficient to complete our benchmarks. We also design multiple test sets to systematically study the generalization of visual learning models, where we vary the overlap of the HOI concepts between the training and test sets of few-shot instances, from partial to no overlaps. Bongard-HOI presents a substantial challenge to today's visual recognition models. The state-of-the-art HOI detection model achieves only 62% accuracy on few-shot binary prediction while even amateur human testers on MTurk have 91% accuracy. With the Bongard-HOI benchmark, we hope to further advance research efforts in visual reasoning, especially in holistic perception-reasoning systems and better representation learning.

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

nvlabs/bongard-hoi officialmentioned in papermentioned on GitHubpytorchNOASSERTION report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

BenchmarkingFew-Shot Image ClassificationFew-Shot LearningHuman-Object Interaction DetectionNovel ConceptsRepresentation LearningVisual Reasoning

Datasets

Introduced by this paper, per the archive.

Bongard-HOI

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Few-Shot Image Classification Bongard-HOI Human (Amateur) Avg. Accuracy 91.42 #1 of 9 Archive leaderboard report
Few-Shot Image Classification Bongard-HOI Meta-Baseline (ImagNet_R50) Avg. Accuracy 55.82 #6 of 9 Archive leaderboard report
Few-Shot Image Classification Bongard-HOI Meta-Baseline (MoCov2_R50) Avg. Accuracy 54.30 #7 of 9 Archive leaderboard report
Few-Shot Image Classification Bongard-HOI Meta-Baseline (Scratch_R50) Avg. Accuracy 54.23 #8 of 9 Archive leaderboard report
Few-Shot Image Classification Bongard-HOI ANIL (ImageNet_R50) Avg. Accuracy 49.74 #9 of 9 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections