Papers › Actor-agnostic Multi-label Action Recognition with Multi-modal Query
Actor-agnostic Multi-label Action Recognition with Multi-modal Query
Anindya Mondal, Sauradip Nag, Joaquin M Prada, Xiatian Zhu, Anjan Dutta
Existing action recognition methods are typically actor-specific due to the intrinsic topological and apparent differences among the actors. This requires actor-specific pose estimation (e.g., humans vs. animals), leading to cumbersome model design complexity and high maintenance costs. Moreover, they often focus on learning the visual modality alone and single-label classification whilst neglecting other available information sources (e.g., class name text) and the concurrent occurrence of multiple actions. To overcome these limitations, we propose a new approach called 'actor-agnostic multi-modal multi-label action recognition,' which offers a unified solution for various types of actors, including humans and animals. We further formulate a novel Multi-modal Semantic Query Network (MSQNet) model in a transformer-based object detection framework (e.g., DETR), characterized by leveraging visual and textual modalities to represent the action classes better. The elimination of actor-specific model designs is a key advantage, as it removes the need for actor pose estimation altogether. Extensive experiments on five publicly available benchmarks show that our MSQNet consistently outperforms the prior arts of actor-specific alternatives on human and animal single- and multi-label action recognition tasks by up to 50%. Code is made available at https://github.com/mondalanindya/MSQNet.
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
1 archive task tag without a task page not shown.
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Action Recognition | Animal Kingdom | MSQNet | mAP | 73.1 | #3 of 4 | Archive leaderboard | report |
| Action Recognition | Charades | MSQNet | MAP | 47.57 | #1 of 1 | Archive leaderboard | report |
| Action Recognition | HMDB51 | MSQNet | Accuracy | 93.25 | #1 of 1 | Archive leaderboard | report |
| Action Recognition | Hockey | MSQNet | Accuracy | 3.05 | #1 of 1 | Archive leaderboard | report |
| Action Recognition | THUMOS14 | MSQNet | Accuracy | 83.16 | #1 of 1 | Archive leaderboard | report |
| Zero-Shot Action Recognition | Charades | MSQNet | mAP | 35.59 | #1 of 4 | Archive leaderboard | report |
| Zero-Shot Action Recognition | HMDB51 | MSQNet | Accuracy | 69.43 | #29 of 29 | Archive leaderboard | report |
| Zero-Shot Action Recognition | THUMOS' 14 | MSQNet | Accuracy | 75.33 | #1 of 1 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Methods
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections