Browse State-of-the-Art › Temporal Sentence Grounding
Temporal Sentence Grounding
16 papers with code · 2 benchmarks · 1 dataset archive 2025-07-28
Temporal sentence grounding (TSG) aims to locate a specific moment from an untrimmed video with a given natural language query. For this task, different levels of supervision are used. 1) Weak supervision: video-level action category set; 2) Semi-weak supervision: video-level action category set, and action annotations at several timestamps; 3) Full supervision: Action category and action interval annotations of all actions in untrimmed videos.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
2 leaderboard tables shown for this task, 2 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| Charades-STA (13 rows) | DeCafNet | DeCafNet: Delegate and Conquer for Efficient Temporal Grounding in... | code | — | Compare |
| Ego4D-Goalstep (3 rows) | DeCafNet-100% | DeCafNet: Delegate and Conquer for Efficient Temporal Grounding in... | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
1 dataset whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
16 shown of 16 papers with code (43 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
10 Sep 2021 2 repositories listedInstead, from a perspective on temporal grounding as a metric-learning problem, we present a Mutual Matching Network (MMN), to directly model the similarity between language queries and video moments in a joint…
-
22 May 2025 1 repository listedThe approach taken by existing methods of dividing video into clips and processing each clip via a full-scale expert encoder is challenging to scale due to prohibitive computational costs of processing a large number of…
-
4 Oct 2024 1 repository listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)Video Large Language Models (Video-LLMs) have demonstrated remarkable capabilities in coarse-grained video understanding, however, they struggle with fine-grained temporal grounding.
-
13 Sep 2024 1 repository listedIn this paper, we address a challenging task, synchronous motion captioning, that aim to generate a language description synchronized with human motion sequences.
-
31 May 2024 1 repository listedTo tackle this limitation, we present a Region-Guided TRansformer (RGTR) for temporal sentence grounding, which diversifies moment queries to eliminate overlapped and redundant predictions.
-
15 Jan 2024 1 repository listed Syntology ran 7 of 7 samples · 0 unverified · 7 pointer-only (licence)Temporal Sentence Grounding in Video (TSGV) is troubled by dataset bias issue, which is caused by the uneven temporal distribution of the target moments for samples with similar semantic components in input videos or…
-
27 Dec 2023 1 repository listedIn the weakly supervised temporal video grounding study, previous methods use predetermined single Gaussian proposals which lack the ability to express diverse events described by the sentence query.
-
30 Nov 2023 1 repository listed Syntology ran 8 of 11 samples · 3 unverified · 11 pointer-only (licence)This separate design allows the model to focus on desirable regions, enabling precise refinement of moment predictions.
-
26 Oct 2023 1 repository listedCompared to traditional benchmarks on which this task is evaluated, these datasets offer finer-grained sentences to ground in notably longer videos.
-
14 Aug 2023 1 repository listedThe goal of TSGSV is to evaluate the relevance between a video stream and a given sentence query.
-
8 Aug 2023 1 repository listed Syntology ran 13 of 17 samples · 4 unverified · 17 pointer-only (licence)Under this setup, we propose a Dynamic Gaussian prior based Grounding framework with Glance annotation (D3G), which consists of a Semantic Alignment Group Contrastive Learning module (SA-GCL) and a Dynamic Gaussian…
-
7 Aug 2023 1 repository listedTo tackle this problem, we propose a novel efficient multi-teacher model (EMTM) based on knowledge distillation to transfer diverse knowledge from both heterogeneous and isomorphic networks.
-
1 Jan 2022 1 repository listedMoreover, they train their model to distinguish positive visual-language pairs from negative ones randomly collected from other videos, ignoring the highly confusing video segments within the same video.
-
1 Sep 2020 1 repository listed Syntology ran 0 of 5 samples · 5 unverifiedIn this paper, we present a series of experiments assessing how well the benchmark results reflect the true progress in solving the moment retrieval task.
-
29 Apr 2020 1 repository listed Syntology ran 1 of 8 samples · 7 unverifiedGiven an untrimmed video and a text query, natural language video localization (NLVL) is to locate a matching span from the video that semantically corresponds to the query.
-
31 Oct 2019 1 repository listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)Temporal sentence grounding in videos aims to detect and localize one target video segment, which semantically corresponds to a given sentence.
Syntology lines on 7 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections