Papers › Scaling Open-Vocabulary Action Detection

Scaling Open-Vocabulary Action Detection

4 Apr 2025arXiv:2504.03096archive 2025-07-28

Zhen Hao Sia, Yogesh Singh Rawat

In this work, we focus on scaling open-vocabulary action detection. Existing approaches for action detection are predominantly limited to closed-set scenarios and rely on complex, parameter-heavy architectures. Extending these models to the open-vocabulary setting poses two key challenges: (1) the lack of large-scale datasets with many action classes for robust training, and (2) parameter-heavy adaptations to a pretrained vision-language contrastive model to convert it for detection, risking overfitting the additional non-pretrained parameters to base action classes. Firstly, we introduce an encoder-only multimodal model for video action detection, reducing the reliance on parameter-heavy additions for video action detection. Secondly, we introduce a simple weakly supervised training strategy to exploit an existing closed-set action detection dataset for pretraining. Finally, we depart from the ill-posed base-to-novel benchmark used by prior works in open-vocabulary action detection and devise a new benchmark to evaluate on existing closed-set action detection datasets without ever using them for training, showing novel results to serve as baselines for future work.

PaperPDFCode

Code

siatheindochinese/sia_act_placeholder officialmentioned on GitHub report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Action DetectionMultiple Action DetectionOpen Vocabulary Action DetectionSpatio-Temporal Action LocalizationVideo Action DetectionZero-Shot Action Detection

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Action Detection J-HMDB SiA Frame-mAP 0.5 88.5 #1 of 18 Archive leaderboard report
Action Detection MultiSports SiA Frame-mAP 0.5 28.8 #2 of 2 Archive leaderboard report
Action Detection UCF101-24 SiA Frame-mAP 0.5 88.5 #2 of 19 Archive leaderboard report
Open Vocabulary Action Detection JHMDB SiA val mAP 57.1 #1 of 1 Archive leaderboard report
Open Vocabulary Action Detection MultiSports SiA val mAP 1.3 #1 of 1 Archive leaderboard report
Open Vocabulary Action Detection UCF101-24 SiA val mAP 42.6 #1 of 1 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

BASEFocus

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections