Papers › Actor-agnostic Multi-label Action Recognition with Multi-modal Query

Actor-agnostic Multi-label Action Recognition with Multi-modal Query

20 Jul 2023arXiv:2307.10763archive 2025-07-28

Anindya Mondal, Sauradip Nag, Joaquin M Prada, Xiatian Zhu, Anjan Dutta

Existing action recognition methods are typically actor-specific due to the intrinsic topological and apparent differences among the actors. This requires actor-specific pose estimation (e.g., humans vs. animals), leading to cumbersome model design complexity and high maintenance costs. Moreover, they often focus on learning the visual modality alone and single-label classification whilst neglecting other available information sources (e.g., class name text) and the concurrent occurrence of multiple actions. To overcome these limitations, we propose a new approach called 'actor-agnostic multi-modal multi-label action recognition,' which offers a unified solution for various types of actors, including humans and animals. We further formulate a novel Multi-modal Semantic Query Network (MSQNet) model in a transformer-based object detection framework (e.g., DETR), characterized by leveraging visual and textual modalities to represent the action classes better. The elimination of actor-specific model designs is a key advantage, as it removes the need for actor pose estimation altogether. Extensive experiments on five publicly available benchmarks show that our MSQNet consistently outperforms the prior arts of actor-specific alternatives on human and animal single- and multi-label action recognition tasks by up to 50%. Code is made available at https://github.com/mondalanindya/MSQNet.

PaperPDFCode

Code

mondalanindya/msqnet officialmentioned in papermentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Action ClassificationAction RecognitionAction Recognition In VideosAnimal Action RecognitionZero-Shot Action Recognition

1 archive task tag without a task page not shown.

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Action Recognition Animal Kingdom MSQNet mAP 73.1 #3 of 4 Archive leaderboard report
Action Recognition Charades MSQNet MAP 47.57 #1 of 1 Archive leaderboard report
Action Recognition HMDB51 MSQNet Accuracy 93.25 #1 of 1 Archive leaderboard report
Action Recognition Hockey MSQNet Accuracy 3.05 #1 of 1 Archive leaderboard report
Action Recognition THUMOS14 MSQNet Accuracy 83.16 #1 of 1 Archive leaderboard report
Zero-Shot Action Recognition Charades MSQNet mAP 35.59 #1 of 4 Archive leaderboard report
Zero-Shot Action Recognition HMDB51 MSQNet Accuracy 69.43 #29 of 29 Archive leaderboard report
Zero-Shot Action Recognition THUMOS' 14 MSQNet Accuracy 75.33 #1 of 1 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

AttentionDense ConnectionsFocusLayer NormalizationLinear LayerMulti-Head AttentionResidual ConnectionSoftmaxVision Transformer

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections