Browse State-of-the-Art › Action Segmentation
Action Segmentation
99 papers with code · 9 benchmarks · 18 datasets archive 2025-07-28
Action Segmentation is a challenging problem in high-level video understanding. In its simplest form, Action Segmentation aims to segment a temporally untrimmed video by time and label each segmented part with one of pre-defined action labels. The results of Action Segmentation can be further used as input to various applications, such as video-to-text and action localization.
Source: TricorNet: A Hybrid Temporal Convolutional and Recurrent Network for Video Action Segmentation
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
9 leaderboard tables shown for this task, 9 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
18 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
4 subtasks in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 99 papers with code (219 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
16 Nov 2016 5 repositories listed Syntology ran 2 of 21 samples · 19 unverifiedThe ability to identify and temporally segment fine-grained human actions throughout a video is crucial for robotics, surveillance, education, and beyond.
-
13 Dec 2019 4 repositories listed Syntology ran 3 of 5 samples · 2 unverifiedAnnotating videos is cumbersome, expensive and not scalable.
-
20 Mar 2024 3 repositories listedWe found that domain experts prefer our system and find it more informative than purely neural approaches to AQA in diving.
-
19 Oct 2022 3 repositories listed Syntology ran 2 of 4 samples · 2 unverifiedTemporal action segmentation (TAS) in videos aims at densely identifying video frames in minutes-long videos with multiple action classes.
-
7 Apr 2024 2 repositories listed Syntology ran 4 of 4 samples · 0 unverified · 4 pointer-only (licence)Action quality assessment (AQA) has become an emerging topic since it can be extensively applied in numerous scenarios.
-
1 Sep 2022 2 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)This paper introduces a unified framework for video action segmentation via sequence to sequence (seq2seq) translation in a fully and timestamp supervised setup.
-
14 Jun 2022 2 repositories listed Syntology ran 2 of 3 samples · 1 unverified · 3 pointer-only (licence)Our search scheme exploits both global search to find the coarse combinations and local search to get the refined receptive field combinations further.
-
28 Sep 2021 2 repositories listedWe present VideoCLIP, a contrastive approach to pre-train a unified model for zero-shot video and text understanding, without using any labels on downstream tasks.
-
4 Jan 2021 2 repositories listed Syntology ran 3 of 3 samples · 0 unverified · 3 pointer-only (licence)Our search scheme exploits both global search to find the coarse combinations and local search to get the refined receptive field combination patterns further.
-
15 Feb 2020 2 repositories listed Syntology ran 1 of 4 samples · 3 unverified · 1 pointer-only (licence)However, most of the existing multimodal models are pre-trained for understanding tasks, leading to a pretrain-finetune discrepancy for generation tasks.
-
8 Apr 2019 2 repositories listedThe task of temporally detecting and segmenting actions in untrimmed videos has seen an increased attention recently.
-
5 Mar 2019 2 repositories listedTemporally locating and classifying action segments in long untrimmed videos is of particular interest to many applications like surveillance and robotics.
-
11 Jun 2025 1 repository listed Syntology ran 4 of 11 samples · 7 unverifiedIn this work, we pioneer textual reference-guided human action segmentation in multi-person settings, where a textual description specifies the target person for segmentation.
-
2 Jun 2025 1 repository listedThe kitchen is an excellent environment for assessing human motor and cognitive function, as many complex actions are naturally exhibited in kitchens from chopping to cleaning.
-
24 Mar 2025 1 repository listedIn temporal action segmentation approaches, we identified a bi-level learning bias.
-
1 Jan 2025 1 repository listedSubsequently, locally correlated tokens are delivered to the Inter-frame Temporal Mamba module, which integrates long-term point features across the entire video with linear complexity.
-
23 Dec 2024 1 repository listedIn this work, we address unsupervised temporal action segmentation, which segments a set of long, untrimmed videos into semantically meaningful segments that are consistent across videos.
-
2 Nov 2024 1 repository listedTemporal context plays a significant role in temporal action segmentation.
-
31 Oct 2024 1 repository listedRegarding action relationships, the Action Relationships Supervision (ARS) module enhances the discrimination across action classes through contrastive learning of single-class action-text pairs and models the semantic…
-
29 Oct 2024 1 repository listedMultimodal systems have great potential to assist humans in procedural activities, where people follow instructions to achieve their goals.
-
8 Oct 2024 1 repository listedTo address these limitations, we propose a novel method named Language-assisted Human Part Motion Representation Learning (LPL), which contains a Disentangled Part Motion Encoder (DPE) to extract dual-level (i.
-
30 Sep 2024 1 repository listedFor the task of temporal action segmentation, existing works commonly treat it as a frame-wise classification problem.
-
13 Sep 2024 1 repository listedIn this paper, we address a challenging task, synchronous motion captioning, that aim to generate a language description synchronized with human motion sequences.
-
29 Aug 2024 1 repository listedIn the experimental results, we validated the usefulness of 3D pose features as input and the fine-grained dataset for the TAS model in figure skating.
-
19 Aug 2024 1 repository listedHowever, state-of-the-art temporal action segmentation methods overlook the long tail and fail to recognize tail actions.
-
23 Jul 2024 1 repository listedAction segmentation of behavioral videos is the process of labeling each frame as belonging to one or more discrete classes, and is a crucial component of many studies that investigate animal behavior.
-
25 May 2024 1 repository listedAlthough the performance of Temporal Action Segmentation (TAS) has improved in recent years, achieving promising results often comes with a high computational cost due to dense inputs, complex model structures, and…
-
6 May 2024 1 repository listedTo address these limitations, the snippet-aware Transformer with multiple action element (ME-ST) is proposed to enhance the discrimination and segmentation among actions, which leverages intrasnippet attention along…
-
1 Apr 2024 1 repository listedWe evaluate our segmentation approach and unsupervised learning pipeline on the Breakfast, 50-Salads, YouTube Instructions and Desktop Assembly datasets, yielding state-of-the-art results for the unsupervised video…
-
28 Mar 2024 1 repository listedWeakly-supervised action segmentation is a task of learning to partition a long video into several action segments, where training videos are only accompanied by transcripts (ordered list of actions).
Syntology lines on 9 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections