Browse State-of-the-Art › Action Understanding
Action Understanding
35 papers with code · 1 benchmark · 4 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
1 leaderboard table shown for this task, 1 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| Win-Fail Action Understanding (1 row) | 2DCNN+TRN | Win-Fail Action Recognition | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
4 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 35 papers with code (88 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
8 Aug 2023 2 repositories listedMoreover, combining these two paradigms in a naive manner leaves the synergy between them untapped and can lead to interference in training.
-
LLaVA-Pose: Enhancing Human Pose and Action Understanding via Keypoint-Integrated Instruction Tuning26 Jun 2025 1 repository listedWe fine-tune the LLaVA-1.
-
11 Apr 2025 1 repository listed Syntology ran 9 of 12 samples · 3 unverified · 12 pointer-only (licence)Analyzing Fast, Frequent, and Fine-grained (F³) events presents a significant challenge in video analytics and multi-modal LLMs.
-
24 Mar 2025 1 repository listedThe recent development of multi-modal large language models (MLLMs) is a promising candidate for a wide range of action understanding tasks.
-
2 Jan 2025 1 repository listedExperiments show that SeFAR achieves state-of-the-art performance on two FAR datasets, FineGym and FineDiving, across various data scopes.
-
31 Oct 2024 1 repository listedRegarding action relationships, the Action Relationships Supervision (ARS) module enhances the discrimination across action classes through contrastive learning of single-class action-text pairs and models the semantic…
-
13 Jun 2024 1 repository listed Syntology ran 9 of 9 samples · 0 unverifiedTo facilitate research on egocentric and exocentric full-body action understanding, we construct benchmarks on a suite of tasks (i.
-
11 Jun 2024 1 repository listed Syntology ran 10 of 16 samples · 6 unverifiedWith approximately 285 hours of surgical videos, OphNet is about 20 times larger than the largest existing surgical workflow analysis benchmark.
-
5 Jun 2024 1 repository listedFollowing the taxonomy of context-based, generative learning, and contrastive learning approaches, we make a thorough review and benchmark of existing works and shed light on the future possible directions.
-
FineParser: A Fine-grained Spatio-temporal Action Parser for Human-centric Action Quality Assessment11 May 2024 1 repository listed Syntology ran 4 of 5 samples · 1 unverified · 5 pointer-only (licence)We argue that a fine-grained understanding of actions requires the model to perceive and parse actions in both time and space, which is also the key to the credibility and interpretability of the AQA technique.
-
3 Jan 2024 1 repository listedReasoning over sports videos for question answering is an important task with numerous applications, such as player training and information retrieval.
-
1 Jan 2024 1 repository listedFine-grained action analysis in multi-person sports is complex due to athletes' quick movements and intense physical confrontations which result in severe visual obstructions in most scenes.
-
25 Dec 2023 1 repository listedA comprehensive understanding of videos is inseparable from describing the action with its contextual action-object interactions.
-
6 Nov 2023 1 repository listedIn this manner, our framework is able to learn the unified representations of uni-modal or multi-modal skeleton input, which is flexible to different kinds of modality input for robust action understanding in practical…
-
15 Aug 2023 1 repository listed Syntology ran 0 of 1 samples · 1 unverifiedBased on this idea, we present Memory-and-Anticipation Transformer (MAT), a memory-anticipation-based approach, to address the online action detection and anticipation tasks.
-
18 May 2023 1 repository listed Syntology ran 1 of 2 samples · 1 unverified · 2 pointer-only (licence)Action knowledge involves the understanding of textual, visual, and temporal aspects of actions.
-
1 Oct 2022 1 repository listedThis paves the way for a systematic way of evaluating embodied AI agents that understand grounded actions.
-
Weakly-Supervised Temporal Action Detection for Fine-Grained Videos with Hierarchical Atomic Actions24 Jul 2022 1 repository listedAction understanding has evolved into the era of fine granularity, as most human behaviors in real life have only minor differences.
-
19 Jul 2022 1 repository listedAction Quality Assessment(AQA) is important for action understanding and resolving the task poses unique challenges due to subtle visual differences.
-
28 Apr 2022 1 repository listedIn videos that contain actions performed unintentionally, agents do not achieve their desired goals.
-
26 Mar 2022 1 repository listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)The generated text prompts are paired with corresponding video clips, and together co-train the text encoder and the video encoder via a contrastive approach.
-
28 Feb 2022 1 repository listedTo that end, we propose to learn exercise-oriented image and video representations from unlabeled samples such that a small dataset annotated by experts suffices for supervised error detection.
-
22 Nov 2021 1 repository listedFor human action understanding, a popular research direction is to analyze short video clips with unambiguous semantic content, such as jumping and drinking.
-
3 Sep 2021 1 repository listedThis leads to poor accuracy when downstream tasks, such as action recognition, depend on pose.
-
15 Aug 2021 1 repository listedFine-grained action recognition is attracting increasing attention due to the emerging demand of specific action understanding in real-world applications, whereas the data of rare fine-grained categories is very limited.
-
21 Jun 2021 1 repository listed Syntology ran 5 of 7 samples · 2 unverified · 7 pointer-only (licence)Hand modeling is critical for immersive VR/AR, action understanding, or human healthcare.
-
11 May 2021 1 repository listed Syntology ran 2 of 16 samples · 14 unverifiedHowever, there remains a lack of studies that extend action composition and leverage multiple viewpoints and multiple modalities of data for representation learning.
-
15 Feb 2021 1 repository listedWe introduce a first of its kind paired win-fail action understanding dataset with samples from the following domains: "General Stunts," "Internet Wins-Fails," "Trick Shots," and "Party Games."
-
14 Dec 2020 1 repository listedThe main reason is that large number of nodes (i.
-
13 Oct 2020 1 repository listedMany believe that the successes of deep learning on image understanding problems can be replicated in the realm of video understanding.
Syntology lines on 9 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections