Papers › Temporal Reasoning Graph for Activity Recognition

Temporal Reasoning Graph for Activity Recognition

27 Aug 2019arXiv:1908.09995archive 2025-07-28

Jingran Zhang, Fumin Shen, Xing Xu, Heng Tao Shen

Despite great success has been achieved in activity analysis, it still has many challenges. Most existing work in activity recognition pay more attention to design efficient architecture or video sampling strategy. However, due to the property of fine-grained action and long term structure in video, activity recognition is expected to reason temporal relation between video sequences. In this paper, we propose an efficient temporal reasoning graph (TRG) to simultaneously capture the appearance features and temporal relation between video sequences at multiple time scales. Specifically, we construct learnable temporal relation graphs to explore temporal relation on the multi-scale range. Additionally, to facilitate multi-scale temporal relation extraction, we design a multi-head temporal adjacent matrix to represent multi-kinds of temporal relations. Eventually, a multi-head temporal relation aggregator is proposed to extract the semantic meaning of those features convolving through the graphs. Extensive experiments are performed on widely-used large-scale datasets, such as Something-Something and Charades, and the results show that our model can achieve state-of-the-art performance. Further analysis shows that temporal relation reasoning with our TRG can extract discriminative features for activity recognition.

PaperPDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Action RecognitionActivity RecognitionRelation ExtractionTemporal Relation Extraction

1 archive task tag without a task page not shown.

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Action Recognition Something-Something V1 TRG (Inception-V3) Top 1 Accuracy 49.7 #55 of 74 Archive leaderboard report
Action Recognition Something-Something V1 TRG (ResNet-50) Top 1 Accuracy 49.5 #56 of 74 Archive leaderboard report
Action Recognition Something-Something V1 TRG (ResNet-50) Top 5 Accuracy 86.1 #56 of 74 Archive leaderboard report
Action Recognition Something-Something V2 TRG (ResNet-50) Top-1 Accuracy 62.2 #105 of 123 Archive leaderboard report
Action Recognition Something-Something V2 TRG (ResNet-50) Top-5 Accuracy 90.3 #105 of 123 Archive leaderboard report
Action Recognition Something-Something V2 TRG (Inception-V3) Top-1 Accuracy 61.3 #109 of 123 Archive leaderboard report
Action Recognition Something-Something V2 TRG (Inception-V3) Top-5 Accuracy 91.4 #109 of 123 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

1x1 ConvolutionAverage PoolingBatch NormalizationBottleneck Residual BlockConvolutionGlobal Average PoolingKaiming InitializationMax PoolingReLUResidual BlockResidual Connection

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections