Browse State-of-the-Art › Natural Language Moment Retrieval
Natural Language Moment Retrieval
21 papers with code · 4 benchmarks · 5 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
4 leaderboard tables shown for this task, 4 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| TACoS (13 rows) | SG-DETR (w/ PT) | Saliency-Guided DETR for Moment Retrieval and Highlight Detection | code | — | Compare |
| ActivityNet Captions (8 rows) | GVL (paragraph-level) | Learning Grounded Vision-Language Representation for Versatile... | code | Syntology ran 1 of 12 samples · 11 unverified | Compare |
| MAD (8 rows) | ReVisionLLM | ReVisionLLM: Recursive Vision-Language Model for Temporal... | code | — | Compare |
| DiDeMo (1 row) | VLG-Net | VLG-Net: Video-Language Graph Matching Network for Video Grounding | code | Syntology ran 1 of 1 samples · 0 unverified | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
5 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
21 shown of 21 papers with code (22 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
11 Dec 2023 2 repositories listedAdapting existing short video (5-30 seconds) grounding methods to this problem yields poor performance.
-
15 Nov 2023 2 repositories listed Syntology ran 2 of 11 samples · 9 unverified · 11 pointer-only (licence)Dummy tokens conditioned by text query take portions of the attention weights, preventing irrelevant video clips from being represented by the text query.
-
22 May 2025 1 repository listedThe approach taken by existing methods of dividing video into clips and processing each clip via a full-scale expert encoder is challenging to scale due to prohibitive computational costs of processing a large number of…
-
18 Jan 2025 1 repository listedExisting models usually first use contrastive learning methods to align video and text features, then fuse and extract multimodal information, and finally use a Transformer Decoder to decode multimodal information.
-
18 Dec 2024 1 repository listed Syntology ran 2 of 13 samples · 11 unverified · 13 pointer-only (licence)For short-moment retrieval, FlashVTG increases mAP to 125% of previous SOTA performance.
-
22 Nov 2024 1 repository listedWe propose ReVisionLLM, a recursive vision-language model designed to locate events in hour-long videos.
-
21 Nov 2024 1 repository listedHowever, long video processing and precise moment retrieval remain challenging due to LLMs' limited context size and coarse frame extraction.
-
2 Oct 2024 1 repository listedCombined with the introduced Saliency-Guided Cross Attention mechanism and a hybrid DETR architecture, our approach significantly enhances performance in both moment retrieval and highlight detection tasks.
-
21 Jul 2024 1 repository listed Syntology ran 5 of 11 samples · 6 unverifiedThrough a feasibility study, we demonstrate that LLM encoders effectively refine inter-concept relations in multimodal embeddings, even without being trained on textual embeddings.
-
26 Jun 2024 1 repository listed Syntology ran 0 of 1 samples · 1 unverifiedWe achieve a new state-of-the-art in moment retrieval on the widely used benchmarks Charades-STA, QVHighlights, and ActivityNet Captions.
-
7 Apr 2024 1 repository listedTemporal Action Detection (TAD) focuses on detecting pre-defined actions, while Moment Retrieval (MR) aims to identify the events described by open-ended natural language within untrimmed videos.
-
30 Nov 2023 1 repository listed Syntology ran 8 of 11 samples · 3 unverified · 11 pointer-only (licence)This separate design allows the model to focus on desirable regions, enabling precise refinement of moment predictions.
-
28 Nov 2023 1 repository listed Syntology ran 9 of 12 samples · 3 unverifiedVideo Moment Retrieval (MR) and Highlight Detection (HD) have attracted significant attention due to the growing demand for video analysis.
-
21 Aug 2023 1 repository listedWhile large-scale image-text pretrained models such as CLIP have been used for multiple video-level tasks on trimmed videos, their use for temporal localization in untrimmed videos is still a relatively unexplored task.
-
31 Jul 2023 1 repository listed Syntology ran 10 of 16 samples · 6 unverifiedMost methods in this direction develop taskspecific models that are trained with type-specific labels, such as moment retrieval (time interval) and highlight detection (worthiness curve), which limits their abilities to…
-
5 Jun 2023 1 repository listedVideo moment retrieval (VMR) identifies a specific moment in an untrimmed video for a given natural language query.
-
11 Mar 2023 1 repository listed Syntology ran 1 of 12 samples · 11 unverifiedOur framework is easily extensible to tasks covering visually-grounded language understanding and generation.
-
26 Feb 2023 1 repository listedIn this paper, we propose a method for improving the performance of natural language grounding in long videos by identifying and pruning out non-describable windows.
-
1 Dec 2021 1 repository listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)The recent and increasing interest in video-language research has driven the development of large-scale datasets that enable data-intensive machine learning techniques.
-
19 Nov 2020 1 repository listed Syntology ran 1 of 1 samples · 0 unverifiedGrounding language queries in videos aims at identifying the time interval (or moment) semantically relevant to a language query.
-
7 Apr 2020 1 repository listedThe key idea of this paper is to use the distances between the frame within the ground truth and the starting (ending) frame as dense supervisions to improve the video grounding accuracy.
Syntology lines on 10 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections