Papers › Cross-Enhancement Transformer for Action Segmentation

Cross-Enhancement Transformer for Action Segmentation

19 May 2022arXiv:2205.09445archive 2025-07-28

Jiahui Wang, Zhenyou Wang, Shanna Zhuang, Hui Wang

Temporal convolutions have been the paradigm of choice in action segmentation, which enhances long-term receptive fields by increasing convolution layers. However, high layers cause the loss of local information necessary for frame recognition. To solve the above problem, a novel encoder-decoder structure is proposed in this paper, called Cross-Enhancement Transformer. Our approach can be effective learning of temporal structure representation with interactive self-attention mechanism. Concatenated each layer convolutional feature maps in encoder with a set of features in decoder produced via self-attention. Therefore, local and global information are used in a series of frame actions simultaneously. In addition, a new loss function is proposed to enhance the training process that penalizes over-segmentation errors. Experiments show that our framework performs state-of-the-art on three challenging datasets: 50Salads, Georgia Tech Egocentric Activities and the Breakfast dataset.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

Wangjhdeveloper/CETNet officialmentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Action SegmentationDecoderSegmentation

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Action Segmentation 50 Salads CETNet Acc 86.9 #10 of 28 Archive leaderboard report
Action Segmentation 50 Salads CETNet Edit 81.7 #10 of 28 Archive leaderboard report
Action Segmentation 50 Salads CETNet F1@10% 87.6 #10 of 28 Archive leaderboard report
Action Segmentation 50 Salads CETNet F1@25% 86.5 #10 of 28 Archive leaderboard report
Action Segmentation 50 Salads CETNet F1@50% 80.1 #10 of 28 Archive leaderboard report
Action Segmentation Breakfast CETNet Acc 74.9 #7 of 37 Archive leaderboard report
Action Segmentation Breakfast CETNet Average F1 71.8 #7 of 37 Archive leaderboard report
Action Segmentation Breakfast CETNet Edit 77.8 #7 of 37 Archive leaderboard report
Action Segmentation Breakfast CETNet F1@10% 79.3 #7 of 37 Archive leaderboard report
Action Segmentation Breakfast CETNet F1@25% 74.3 #7 of 37 Archive leaderboard report
Action Segmentation Breakfast CETNet F1@50% 61.9 #7 of 37 Archive leaderboard report
Action Segmentation GTEA CETNet Acc 80.3 #9 of 28 Archive leaderboard report
Action Segmentation GTEA CETNet Edit 87.9 #9 of 28 Archive leaderboard report
Action Segmentation GTEA CETNet F1@10% 91.8 #9 of 28 Archive leaderboard report
Action Segmentation GTEA CETNet F1@25% 91.2 #9 of 28 Archive leaderboard report
Action Segmentation GTEA CETNet F1@50% 81.3 #9 of 28 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Absolute Position EncodingsAdamAttentionBPEConvolutionDense ConnectionsDropoutHTCNLabel SmoothingLayer NormalizationLinear LayerMulti-Head AttentionPosition-Wise Feed-Forward LayerResidual ConnectionSoftmaxTransformer

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections