Browse State-of-the-Art › Temporal Localization
Temporal Localization
76 papers with code · 0 benchmarks · 4 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
4 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
2 subtasks in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 76 papers with code (153 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
5 May 2017 12 repositories listedFor evaluation, we adopt TaCoS dataset, and build a new dataset for this task on top of Charades by adding sentence temporal annotations, called Charades-STA.
-
21 Nov 2018 3 repositories listedPrevious methods address the problem by considering features from video sliding windows and language queries and learning a subspace to encode their correlation, which ignore rich semantic cues about activities in…
-
14 Dec 2017 3 repositories listedWe propose a weakly supervised temporal action localization algorithm on untrimmed videos using convolutional neural networks.
-
4 Dec 2023 2 repositories listed Syntology ran 7 of 11 samples · 4 unverifiedThis work proposes TimeChat, a time-sensitive multimodal large language model specifically designed for long video understanding.
-
3 Jun 2022 2 repositories listed Syntology ran 3 of 3 samples · 0 unverified · 2 pointer-only (licence)Video-Language Pretraining (VLP), which aims to learn transferable representation to advance a wide range of video-text downstream tasks, has recently received increasing attention.
-
30 Jul 2019 2 repositories listed Syntology ran 0 of 11 samples · 11 unverifiedWe evaluate our approach on two recently proposed datasets for temporal localization of moments in video with natural language (DiDeMo and Charades-STA) extended to our video corpus moment retrieval setting.
-
26 May 2019 2 repositories listedAmong other uses, VERA enables the localization of a shooter from just a few videos that include the sound of gunshots.
-
23 Mar 2018 2 repositories listed Syntology ran 3 of 4 samples · 1 unverified · 4 pointer-only (licence)In this paper, we introduce a novel problem of audio-visual event localization in unconstrained videos.
-
26 Dec 2017 2 repositories listed Syntology ran 0 of 1 samples · 1 unverifiedThis paper presents a new large-scale dataset for recognition and temporal localization of human actions collected from Web videos.
-
19 Dec 2016 2 repositories listedActions are more than just movements and trajectories: we cook to eat and we hold a cup to drink from it.
-
5 Jun 2025 1 repository listedCurrent video-based approaches, while proficient in tracking, lack the sophisticated reasoning capabilities of large language models, limiting their contextual understanding and generalization.
-
30 May 2025 1 repository listed Syntology ran 1 of 1 samples · 0 unverifiedExisting methods for temporal expression either conflate time with text-based numerical values, add a series of dedicated temporal tokens, or regress time using specialized temporal grounding heads.
-
4 May 2025 1 repository listedConducted experiments confirm applicability of the proposed face ReID algorithm that is combining the concepts of face detection, face recognition and passive tracking-by-detection in order to achieve robust and…
-
1 May 2025 1 repository listedTo remedy this, we provide a new video reasoning dataset called MINERVA for modern multimodal models.
-
24 Apr 2025 1 repository listedTo capture the complexity in human activities, DARai is annotated at three levels of hierarchy: (i) high-level activities (L1) that are independent tasks, (ii) lower-level actions (L2) that are patterns shared between…
-
24 Mar 2025 1 repository listedTo bridge this gap, we introduce the Aerial Traffic Atomic Activity Recognition and Segmentation (ATARS) dataset, the first aerial dataset designed for multi-label atomic activity analysis.
-
17 Mar 2025 1 repository listedIn the domain of audio-visual event perception, which focuses on the temporal localization and classification of events across distinct modalities (audio and visual), existing approaches are constrained by the…
-
17 Mar 2025 1 repository listed Syntology ran 3 of 4 samples · 1 unverifiedVideos, with their unique temporal dimension, demand precise grounded understanding, where answers are directly linked to visual, interpretable evidence.
-
17 Mar 2025 1 repository listedFurthermore, we also visualize the process of explicit cooperation and surprisingly find that each LoRA head has certain audio-visual understanding ability.
-
9 Mar 2025 1 repository listedTemporal localization in untrimmed videos, which aims to identify specific timestamps, is crucial for video understanding but remains challenging.
-
28 Feb 2025 1 repository listedMarine ecosystem monitoring via Passive Acoustic Monitoring (PAM) generates vast data, but deep learning often requires precise annotations and short segments.
-
16 Feb 2025 1 repository listedCompared to existing methods leveraging zero-initialized queries, object queries in our TA-STVG, directly generated from a given video-text pair, naturally carry target-specific cues, making them adaptive and better…
-
14 Jan 2025 1 repository listed Syntology ran 3 of 7 samples · 4 unverifiedFurthermore, we propose ST-Align dataset with 4.
-
29 Dec 2024 1 repository listedWith the rise of multimodal large language models, accurately extracting and understanding textual information from video content, referred to as video based optical character recognition (Video OCR), has become a…
-
12 Dec 2024 1 repository listedRecent work has focused on enabling Video LLMs to perform video temporal grounding via next-token prediction of temporal timestamps.
-
27 Nov 2024 1 repository listedRapid development of large language models (LLMs) has significantly advanced multimodal large language models (LMMs), particularly in vision-language tasks.
-
15 Nov 2024 1 repository listed Syntology ran 1 of 11 samples · 10 unverifiedVideo Large Language Models (Vid-LLMs) have made remarkable advancements in comprehending video content for QA dialogue.
-
29 Aug 2024 1 repository listedIn this paper, we propose a Training-Free Video Temporal Grounding (TFVTG) approach that leverages the ability of pre-trained large models.
-
1 Jul 2024 1 repository listed Syntology ran 8 of 12 samples · 4 unverifiedLeveraging Large Language Models' remarkable proficiency in text-based tasks, recent works on Multi-modal LLMs (MLLMs) extend them to other modalities like vision and audio.
-
11 Jun 2024 1 repository listed Syntology ran 10 of 16 samples · 6 unverifiedWith approximately 285 hours of surgical videos, OphNet is about 20 times larger than the largest existing surgical workflow analysis benchmark.
Syntology lines on 11 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections