Browse State-of-the-Art › Egocentric Activity Recognition
Egocentric Activity Recognition
14 papers with code · 2 benchmarks · 4 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
2 leaderboard tables shown for this task, 2 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| EPIC-KITCHENS-55 (7 rows) | DEEP-HAL with ODF+SDF (AssembleNet++) | Self-supervising Action Recognition by Statistical Moment and... | — | — | Compare |
| EGTEA (6 rows) | LaViLa (Finetuned, TimeSformer-L) | Learning Video Representations from Large Language Models | code | Syntology ran 6 of 20 samples · 14 unverified | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
4 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
14 shown of 14 papers with code (26 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
12 Dec 2018 4 repositories listed Syntology ran 0 of 8 samples · 8 unverifiedTo understand the world, we humans constantly need to relate the present to the past, and put events in context.
-
8 Dec 2022 3 repositories listed Syntology ran 6 of 20 samples · 14 unverified · 20 pointer-only (licence)We introduce LaViLa, a new approach to learning video-language representations by leveraging Large Language Models (LLMs).
-
2 May 2019 3 repositories listed Syntology ran 0 of 21 samples · 21 unverifiedSecond, frame-based models perform quite well on action recognition; is pre-training for good image features sufficient or is pre-training for spatio-temporal features valuable for optimal transfer learning?
-
22 May 2019 2 repositories listedOur method is ranked first in the public leaderboard of the EPIC-Kitchens egocentric action anticipation challenge 2019.
-
11 Apr 2023 1 repository listedResearch has shown the complementarity of camera- and inertial-based data for modeling human activities, yet datasets with both egocentric video and inertial-based sensor data remain scarce.
-
26 Jan 2023 1 repository listedHowever, the deficiency of related dataset hinders the development of multi-modal deep learning for egocentric activity recognition.
-
18 Mar 2022 1 repository listed Syntology ran 5 of 5 samples · 0 unverifiedBy utilizing calibrators to embed feature with four different kinds of contexts in parallel, the learnt representation is expected to be more resilient to diverse types of activities.
-
16 Apr 2021 1 repository listedWe introduce an approach for pre-training egocentric video models using large-scale third-person video datasets.
-
8 Nov 2020 1 repository listedIn addition, we model the distribution of gaze fixations using a variational method.
-
22 Aug 2019 1 repository listedWe focus on multi-modal fusion for egocentric action recognition, and propose a novel architecture for multi-modal temporal-binding, i.
-
26 Nov 2018 1 repository listedEgocentric activity recognition is one of the most challenging tasks in video analysis.
-
31 Jul 2018 1 repository listedOur model is built on the observation that egocentric activities are highly characterized by the objects and their locations in the video.
-
15 Nov 2017 1 repository listedThe per-frame (per-segment) extracted features are considered as a set of time series, and inter and intra-time series relations are employed to represent the video descriptors.
-
8 Apr 2017 1 repository listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)Our dataset and experiments can be of interest to communities of 3D hand pose estimation, 6D object pose, and robotics as well as action recognition.
Syntology lines on 5 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections