Browse State-of-the-Art › Action Anticipation
Action Anticipation
49 papers with code · 8 benchmarks · 11 datasets archive 2025-07-28
Next action anticipation is defined as observing 1, ... , T frames and predicting the action that happens after a gap of T_a seconds. It is important to note that a new action starts after T_a seconds that is not seen in the observed frames. Here T_a=1 second.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
8 leaderboard tables shown for this task, 8 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
11 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 49 papers with code (110 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
23 Jun 2020 7 repositories listed Syntology ran 3 of 18 samples · 15 unverifiedThis paper introduces the pipeline to extend the largest dataset in egocentric vision, EPIC-KITCHENS.
-
1 Jun 2020 2 repositories listedFuture prediction, especially in long-range videos, requires reasoning from current and past observations.
-
4 May 2020 2 repositories listedThe experiments show that the proposed architecture is state-of-the-art in the domain of egocentric videos, achieving top performances in the 2019 EPIC-Kitchens egocentric action anticipation challenge.
-
10 Dec 2019 2 repositories listedThe hallucination task is treated as an auxiliary task, which can be used with any other action related task in a multitask learning setting.
-
22 May 2019 2 repositories listedOur method is ranked first in the public leaderboard of the EPIC-Kitchens egocentric action anticipation challenge 2019.
-
8 Apr 2018 2 repositories listedFirst-person vision is gaining interest as it offers a unique viewpoint on people's interaction with objects, their attention, and even intention.
-
11 Jun 2025 1 repository listed Syntology ran 0 of 8 samples · 8 unverified · 8 pointer-only (licence)Finally, we show how self-supervised learning can be applied to robotic planning tasks by post-training a latent action-conditioned world model, V-JEPA 2-AC, using less than 62 hours of unlabeled robot videos from the…
-
24 Apr 2025 1 repository listedTo capture the complexity in human activities, DARai is annotated at three levels of hierarchy: (i) high-level activities (L1) that are independent tasks, (ii) lower-level actions (L2) that are patterns shared between…
-
1 Jan 2025 1 repository listedExtensive experiments on benchmark datasets demonstrate the superiority of the proposed ActionLLM framework, encouraging a promising direction to explore LLMs in the context of action anticipation.
-
1 Jan 2025 1 repository listedLong-term dense action anticipation is very challenging since it requires predicting actions and their durations several minutes into the future based on provided video observations.
-
17 Sep 2024 1 repository listedVideo Language Models (VLMs) are crucial for generalizing across diverse tasks and using language cues to enhance learning.
-
5 Aug 2024 1 repository listedARR decomposes the action anticipation task into action recognition and sequence reasoning tasks, and effectively learns the statistical relationship between actions by next action prediction (NAP).
-
18 Jul 2024 1 repository listedIn this paper, we present our solutions for a spectrum of automation tasks in life-saving intervention procedures within the Trauma THOMPSON (T3) Challenge, encompassing action recognition, action anticipation, and…
-
16 Jul 2024 1 repository listedAs generator, we introduce a Gated Anticipation Network (GTAN) to model both observed and unobserved frames of a video in a mutual representation.
-
2 Jul 2024 1 repository listed Syntology ran 6 of 8 samples · 2 unverifiedAction anticipation is the task of forecasting future activity from a partially observed sequence of events.
-
26 Jun 2024 1 repository listed Syntology ran 11 of 15 samples · 4 unverified · 15 pointer-only (licence)In this report, we present our solutions to the EgoVis Challenges in CVPR 2024, including five tracks in the Ego4D challenge and three tracks in the EPIC-Kitchens challenge.
-
24 Mar 2024 1 repository listed Syntology ran 7 of 8 samples · 1 unverifiedAlong with the videos we record high-quality gaze data and provide detailed multimodal annotations, formulating a playground for modeling the human ability to bridge asynchronous procedural actions from different…
-
6 Dec 2023 1 repository listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)We present Egocentric Action Scene Graphs (EASGs), a new representation for long-form understanding of egocentric videos.
-
31 Oct 2023 1 repository listed Syntology ran 1 of 5 samples · 4 unverifiedTo recognize and predict human-object interactions, we use a Transformer-based neural architecture which allows the "retrieval" of relevant objects for action anticipation at various time scales.
-
31 Jul 2023 1 repository listed Syntology ran 2 of 2 samples · 0 unverifiedWe propose to formulate the LTA task from two perspectives: a bottom-up approach that predicts the next actions autoregressively by modeling temporal dynamics; and a top-down approach that infers the goal of the actor…
-
4 Jul 2023 1 repository listed Syntology ran 1 of 2 samples · 1 unverifiedIn this report, we describe the technical details of our approach for the Ego4D Long-Term Action Anticipation Challenge 2023.
-
28 Jun 2023 1 repository listed Syntology ran 1 of 6 samples · 5 unverifiedWe present Palm, a solution to the Long-Term Action Anticipation (LTA) task utilizing vision-language and large language models.
-
26 Jun 2023 1 repository listedIn this paper, we address the problem of short-term action anticipation, i.
-
22 May 2023 1 repository listedTo this end, we propose a novel approach that applies a guided attention mechanism between the objects, and the spatiotemporal features extracted from video clips, enhancing the motion and contextual information, and…
-
7 Feb 2023 1 repository listedObject affordance is an important concept in hand-object interaction, providing information on action possibilities based on human motor capacity and objects' physical property thus benefiting tasks such as action…
-
25 Nov 2022 1 repository listedOn the EK100 evaluation server, InAViT is the top-performing method on the public leaderboard (at the time of submission) where it outperforms the second-best model by 3.
-
23 Oct 2022 1 repository listedAlthough human action anticipation is a task which is inherently multi-modal, state-of-the-art methods on well known action anticipation datasets leverage this data by applying ensemble methods and averaging scores of…
-
20 Oct 2022 1 repository listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)Action anticipation involves predicting future actions having observed the initial portion of a video.
-
12 Oct 2022 1 repository listedAnticipating future actions in a video is useful for many autonomous and assistive technologies.
-
27 Sep 2022 1 repository listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)However, learning representations from videos can be challenging.
Syntology lines on 12 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections