Papers › Knowing Your Target: Target-Aware Transformer Makes Better Spatio-Temporal Video Grounding

Knowing Your Target: Target-Aware Transformer Makes Better Spatio-Temporal Video Grounding

16 Feb 2025arXiv:2502.11168archive 2025-07-28

Xin Gu, Yaojie Shen, Chenxi Luo, Tiejian Luo, Yan Huang, Yuewei Lin, Heng Fan, Libo Zhang

Transformer has attracted increasing interest in STVG, owing to its end-to-end pipeline and promising result. Existing Transformer-based STVG approaches often leverage a set of object queries, which are initialized simply using zeros and then gradually learn target position information via iterative interactions with multimodal features, for spatial and temporal localization. Despite simplicity, these zero object queries, due to lacking target-specific cues, are hard to learn discriminative target information from interactions with multimodal features in complicated scenarios (\e.g., with distractors or occlusion), resulting in degradation. Addressing this, we introduce a novel Target-Aware Transformer for STVG (TA-STVG), which seeks to adaptively generate object queries via exploring target-specific cues from the given video-text pair, for improving STVG. The key lies in two simple yet effective modules, comprising text-guided temporal sampling (TTS) and attribute-aware spatial activation (ASA), working in a cascade. The former focuses on selecting target-relevant temporal cues from a video utilizing holistic text information, while the latter aims at further exploiting the fine-grained visual attribute information of the object from previous target-aware temporal cues, which is applied for object query initialization. Compared to existing methods leveraging zero-initialized queries, object queries in our TA-STVG, directly generated from a given video-text pair, naturally carry target-specific cues, making them adaptive and better interact with multimodal features for learning more discriminative information to improve STVG. In our experiments on three benchmarks, TA-STVG achieves state-of-the-art performance and significantly outperforms the baseline, validating its efficacy.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

HengLan/TA-STVG officialmentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

AttributeObjectSpatio-Temporal Video GroundingTemporal LocalizationVideo Grounding

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Spatio-Temporal Video Grounding HC-STVG1 TA-STVG m_vIoU 39.1 #1 of 3 Archive leaderboard report
Spatio-Temporal Video Grounding HC-STVG1 TA-STVG vIoU@0.3 63.1 #1 of 3 Archive leaderboard report
Spatio-Temporal Video Grounding HC-STVG1 TA-STVG vIoU@0.5 36.8 #1 of 3 Archive leaderboard report
Spatio-Temporal Video Grounding HC-STVG2 TA-STVG Val m_vIoU 40.2 #1 of 4 Archive leaderboard report
Spatio-Temporal Video Grounding HC-STVG2 TA-STVG Val vIoU@0.3 65.8 #1 of 4 Archive leaderboard report
Spatio-Temporal Video Grounding HC-STVG2 TA-STVG Val vIoU@0.5 36.7 #1 of 4 Archive leaderboard report
Spatio-Temporal Video Grounding VidSTG TA-STVG Declarative m_vIoU 34.4 #1 of 3 Archive leaderboard report
Spatio-Temporal Video Grounding VidSTG TA-STVG Declarative vIoU@0.3 48.2 #1 of 3 Archive leaderboard report
Spatio-Temporal Video Grounding VidSTG TA-STVG Declarative vIoU@0.5 33.5 #1 of 3 Archive leaderboard report
Spatio-Temporal Video Grounding VidSTG TA-STVG Interrogative m_vIoU 29.5 #1 of 3 Archive leaderboard report
Spatio-Temporal Video Grounding VidSTG TA-STVG Interrogative vIoU@0.3 41.5 #1 of 3 Archive leaderboard report
Spatio-Temporal Video Grounding VidSTG TA-STVG Interrogative vIoU@0.5 28.0 #1 of 3 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Absolute Position EncodingsAdamAttentionBPEDense ConnectionsDropoutLabel SmoothingLayer NormalizationLinear LayerMulti-Head AttentionPosition-Wise Feed-Forward LayerResidual ConnectionSETSoftmaxTransformer

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections