Browse State-of-the-Art › Activity Detection
Activity Detection
75 papers with code · 1 benchmark · 12 datasets archive 2025-07-28
Detecting activities in extended videos.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
1 leaderboard table shown for this task, 1 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| AVA-Speech (4 rows) | CNN-BiLSTM_best | A Hybrid CNN-BiLSTM Voice Activity Detector | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
12 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 75 papers with code (380 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
17 Sep 2024 3 repositories listed Syntology ran 2 of 7 samples · 5 unverifiedOur resulting model is the first real-time full-duplex spoken large language model, with a theoretical latency of 160ms, 200ms in practice, and is available at https://github.
-
12 Nov 2022 3 repositories listedEnd-to-end diarization presents an attractive alternative to standard cascaded diarization systems because a single system can handle all aspects of the task at once.
-
23 Feb 2021 3 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)We also report the performance on the ROAD tasks of Slowfast and YOLOv5 detectors, as well as that of the winners of the ICCV2021 ROAD challenge, which highlight the challenges faced by situation awareness in autonomous…
-
4 Nov 2019 3 repositories listed Syntology ran 1 of 1 samples · 0 unverifiedWe introduce pyannote.
-
9 Jun 2019 3 repositories listedIn the end, a posteriori SNR weighted energy difference is applied to the extended pitch segments of the denoised speech signal for detecting voice activity.
-
9 Apr 2018 3 repositories listedIn this paper, we introduce a challenging new dataset, MLB-YouTube, designed for fine-grained activity detection.
-
22 Mar 2017 3 repositories listed Syntology ran 0 of 1 samples · 1 unverifiedWe address the problem of activity detection in continuous, untrimmed video streams.
-
28 Nov 2016 3 repositories listedWe propose a single neural network architecture for two tasks: on-line keyword spotting and voice activity detection.
-
29 Aug 2016 3 repositories listedThis thesis explore different approaches using Convolutional and Recurrent Neural Networks to classify and temporally localize activities on videos, furthermore an implementation to achieve it has been proposed.
-
17 Jul 2023 2 repositories listedWe introduce "ivrit.
-
28 Nov 2021 2 repositories listedIn this paper, we reformulate this task as a single-label prediction problem by encoding the multi-speaker labels with power set.
-
12 Oct 2021 2 repositories listedWe propose a system that combines SAD and a BERT model to perform speaker change detection and speaker role detection (SRD) by chunking ASR transcripts, i.
-
8 Apr 2021 2 repositories listedExperiments on multiple speaker diarization datasets conclude that our model can be used with great success on both voice activity detection and overlapped speech detection.
-
25 Nov 2020 2 repositories listed Syntology ran 3 of 3 samples · 0 unverified · 3 pointer-only (licence)Speech activity detection and speaker diarization are used to extract segments from the videos that contain speech.
-
13 Feb 2020 2 repositories listedWith presence detection, how to collect training data with human presence can have a significant impact on the performance.
-
12 Aug 2019 2 repositories listedIn this paper, we propose "personal VAD", a system to detect the voice activity of a target speaker at the frame level.
-
5 Dec 2017 2 repositories listedIn this paper, we introduce the concept of learning latent super-events from activity videos, and present how it benefits activity detection in continuous videos.
-
3 Jun 2025 1 repository listedIn speaker diarization, traditional clustering-based methods remain widely used in real-world applications.
-
16 May 2025 1 repository listedWe assess the impact of denoising on diarization accuracy and compare various voice activity detection (VAD) models, including self-supervised transformer-based frame-wise VAD models.
-
28 Feb 2025 1 repository listedIn this paper, we investigate the ability of current-generation LLMs to identify text related to environmental activities.
-
17 Feb 2025 1 repository listedTo this end, we developed the VANPY (Voice Analysis in Python) framework for automated pre-processing, feature extraction, and classification of voice data.
-
12 Feb 2025 1 repository listedConclusion: We investigate the Team Time-Out and the StOP?-protocol in the OR, by presenting the first OR dataset with temporal annotations of group activities protocols, and introducing a novel group activity detection…
-
10 Feb 2025 1 repository listedOur simulation results demonstrate that the proposed scheme outperforms state-of-the-art mMIMO-based grant-free massive NOMA schemes with the same access latency.
-
24 Jan 2025 1 repository listedIn this work, we proposed deep learning models for the two tasks of the public MARIO Challenge at MICCAI 2024, designed to detect and forecast changes in nAMD severity with longitudinal retinal OCT.
-
19 Dec 2024 1 repository listedWe address the challenge of WiFi-based temporal activity detection and propose an efficient Dual Pyramid Network that integrates Temporal Signal Semantic Encoders and Local Sensitive Response Encoders.
-
1 Dec 2024 1 repository listedThis work introduces the first framework for reconstructing surgical dialogue from unstructured real-world recordings, which is crucial for characterizing teaching tasks.
-
15 Oct 2024 1 repository listedTo facilitate natural and intuitive interactions with diverse user groups in real-world settings, social robots must be capable of addressing the varying requirements and expectations of these groups while adapting…
-
6 Jun 2024 1 repository listedTwo existing SGS systems are evaluated on the corpus and compared against a baseline X-vector transfer learning strategy, trained on the development subset.
-
30 Jan 2024 1 repository listedThis problem is even more severe in cell-free networks as there are many of these parameters to be acquired.
-
30 Jan 2024 1 repository listedThe results show that our system improves the state-of-the-art on the AMI headset mix, using no oracle information and under full evaluation (no collar and including overlapped speech).
Syntology lines on 5 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections