Papers › Building a Video-and-Language Dataset with Human Actions for Multimodal Logical Inference

Building a Video-and-Language Dataset with Human Actions for Multimodal Logical Inference

27 Jun 2021ACL (mmsr, IWCS) 2021 6arXiv:2106.14137archive 2025-07-28

Riko Suzuki, Hitomi Yanaka, Koji Mineshima, Daisuke Bekki

This paper introduces a new video-and-language dataset with human actions for multimodal logical inference, which focuses on intentional and aspectual expressions that describe dynamic human actions. The dataset consists of 200 videos, 5,554 action labels, and 1,942 action triplets of the form <subject, predicate, object> that can be translated into logical semantic representations. The dataset is expected to be useful for evaluating multimodal inference systems between videos and semantically complicated sentences including negation and quantification.

PaperPDFConference PDFCode

Code

rikos3/HumanActions officialmentioned in paper report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Negation

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections