Papers › Scaling Open-Vocabulary Action Detection
Scaling Open-Vocabulary Action Detection
Zhen Hao Sia, Yogesh Singh Rawat
In this work, we focus on scaling open-vocabulary action detection. Existing approaches for action detection are predominantly limited to closed-set scenarios and rely on complex, parameter-heavy architectures. Extending these models to the open-vocabulary setting poses two key challenges: (1) the lack of large-scale datasets with many action classes for robust training, and (2) parameter-heavy adaptations to a pretrained vision-language contrastive model to convert it for detection, risking overfitting the additional non-pretrained parameters to base action classes. Firstly, we introduce an encoder-only multimodal model for video action detection, reducing the reliance on parameter-heavy additions for video action detection. Secondly, we introduce a simple weakly supervised training strategy to exploit an existing closed-set action detection dataset for pretraining. Finally, we depart from the ill-posed base-to-novel benchmark used by prior works in open-vocabulary action detection and devise a new benchmark to evaluate on existing closed-set action detection datasets without ever using them for training, showing novel results to serve as baselines for future work.
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Action Detection | J-HMDB | SiA | Frame-mAP 0.5 | 88.5 | #1 of 18 | Archive leaderboard | report |
| Action Detection | MultiSports | SiA | Frame-mAP 0.5 | 28.8 | #2 of 2 | Archive leaderboard | report |
| Action Detection | UCF101-24 | SiA | Frame-mAP 0.5 | 88.5 | #2 of 19 | Archive leaderboard | report |
| Open Vocabulary Action Detection | JHMDB | SiA | val mAP | 57.1 | #1 of 1 | Archive leaderboard | report |
| Open Vocabulary Action Detection | MultiSports | SiA | val mAP | 1.3 | #1 of 1 | Archive leaderboard | report |
| Open Vocabulary Action Detection | UCF101-24 | SiA | val mAP | 42.6 | #1 of 1 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Methods
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections