Papers › Video Action Transformer Network

Video Action Transformer Network

6 Dec 2018CVPR 2019 6arXiv:1812.02707archive 2025-07-28

Rohit Girdhar, João Carreira, Carl Doersch, Andrew Zisserman

We introduce the Action Transformer model for recognizing and localizing human actions in video clips. We repurpose a Transformer-style architecture to aggregate features from the spatiotemporal context around the person whose actions we are trying to classify. We show that by using high-resolution, person-specific, class-agnostic queries, the model spontaneously learns to track individual people and to pick up on semantic context from the actions of others. Additionally its attention mechanism learns to emphasize hands and faces, which are often crucial to discriminate an action - all without explicit supervision other than boxes and class labels. We train and test our Action Transformer network on the Atomic Visual Actions (AVA) dataset, outperforming the state-of-the-art by a significant margin using only raw RGB frames as input.

PaperPDFConference PDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Action RecognitionRecognizing And Localizing Human Actions

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Action Recognition AVA v2.1 I3D Tx HighRes GFlops 39.6 #6 of 15 Archive leaderboard report
Action Recognition AVA v2.1 I3D Tx HighRes Params (M) 19.3 #6 of 15 Archive leaderboard report
Action Recognition AVA v2.1 I3D Tx HighRes mAP (Val) 27.6 #6 of 15 Archive leaderboard report
Action Recognition AVA v2.1 I3D I3D GFlops 6.5 #10 of 15 Archive leaderboard report
Action Recognition AVA v2.1 I3D I3D Params (M) 16.2 #10 of 15 Archive leaderboard report
Action Recognition AVA v2.1 I3D I3D mAP (Val) 23.4 #10 of 15 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Absolute Position EncodingsAdamAttentionBPEDense ConnectionsDropoutLabel SmoothingLayer NormalizationLinear LayerMulti-Head AttentionPosition-Wise Feed-Forward LayerReLUResidual ConnectionSoftmaxTransformer

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections