Browse State-of-the-Art › Video Saliency Detection
Video Saliency Detection
21 papers with code · 5 benchmarks · 3 datasets archive 2025-07-28
Video Saliency Detection is the process of detecting or predicting the most visually important or attention-grabbing regions in video sequence. This regions are where humans most likely focus their attention when watching the video.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
5 leaderboard tables shown for this task, 5 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| MSU Video Saliency Prediction (14 rows) | ViNet (dave) | ViNet: Pushing the limits of Visual Modality for Audio-Visual... | code | — | Compare |
| DHF1K (3 rows) | ViNet | ViNet: Pushing the limits of Visual Modality for Audio-Visual... | code | — | Compare |
| DIEM (1 row) | AViNet | ViNet: Pushing the limits of Visual Modality for Audio-Visual... | code | — | Compare |
| Hollywood2 (1 row) | ViNet | ViNet: Pushing the limits of Visual Modality for Audio-Visual... | code | — | Compare |
| UCFSports (1 row) | ViNet | ViNet: Pushing the limits of Visual Modality for Audio-Visual... | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
3 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
21 shown of 21 papers with code (36 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
18 Feb 2019 4 repositories listed Syntology ran 0 of 6 samples · 6 unverifiedTo develop robust representations for this challenging task, high-level visual features at multiple spatial scales must be extracted and augmented with contextual information.
-
11 Mar 2020 2 repositories listed Syntology ran 3 of 9 samples · 6 unverifiedWe evaluate our method on the video saliency datasets DHF1K, Hollywood-2 and UCF-Sports, and the image saliency datasets SALICON and MIT300.
-
3 Jul 2019 2 repositories listedThis paper investigates modifying an existing neural network architecture for static saliency prediction using two types of recurrences that integrate information from the temporal domain.
-
23 Sep 2024 1 repository listedThe goal of the participants was to develop a method for predicting accurate saliency maps for the provided set of video sequences.
-
18 Jun 2024 1 repository listedHowever, the existing salient object detection (SOD) works only focus on either static RGB-D images or RGB videos, ignoring the collaborating of RGB-D and video information.
-
5 Dec 2023 1 repository listedIn this work, we present an integrated system for spatiotemporal summarization of 360-degrees videos.
-
19 Sep 2022 1 repository listed Syntology ran 5 of 7 samples · 2 unverified360° video saliency detection is one of the challenging benchmarks for 360° video understanding since non-negligible distortion and discontinuity occur in the projection of any format of 360° videos, and capture-worthy…
-
20 Jun 2022 1 repository listedVideo saliency detection (VSD) aims at fast locating the most attractive objects/things/patterns in a given video clip.
-
9 Jun 2022 1 repository listedWe show that gaze direction and affective representations contribute a prediction to ground-truth correspondence improvement of at least 5% compared to dynamic saliency models without social cues.
-
27 Dec 2021 1 repository listedMoreover, we distill knowledge from these regions to obtain complete new spatial-temporal-audio (STA) fixation prediction (FP) networks, enabling broad applications in cases where video tags are not available.
-
19 Jun 2021 1 repository listedThanks to the rapid advances in the deep learning techniques and the wide availability of large-scale training sets, the performances of video saliency detection models have been improving steadily and significantly.
-
6 Apr 2021 1 repository listedSignificant performance improvement has been achieved for fully-supervised video salient object detection with the pixel-wise labeled training datasets, which are time-consuming and expensive to obtain.
-
11 Dec 2020 1 repository listedWe also explore a variation of ViNet architecture by augmenting audio features into the decoder.
-
2 Oct 2020 1 repository listedWhen the base hierarchical model is empowered with domain-specific modules, performance improves, outperforming state-of-the-art models on three out of five metrics on the DHF1K benchmark and reaching the second-best…
-
7 Aug 2020 1 repository listedIn this way, even though the overall video saliency quality is heavily dependent on its spatial branch, however, the performance of the temporal branch still matter.
-
2 Aug 2020 1 repository listedWith the rapid development of deep learning techniques, image saliency deep models trained solely by spatial information have occasionally achieved detection performance for video data comparable to that of the models…
-
2 Jan 2020 1 repository listedDue to a variety of motions across different frames, it is highly challenging to learn an effective spatiotemporal representation for accurate video saliency prediction (VSP).
-
15 Aug 2019 1 repository listed Syntology ran 0 of 2 samples · 2 unverifiedIt consists of two building blocks: first, the encoder network extracts low-resolution spatiotemporal features from an input clip of several consecutive frames, and then the following prediction network decodes the…
-
7 Apr 2019 1 repository listedIn this work, we proposed an end-to-end dilated inception network (DINet) for visual saliency prediction.
-
1 Sep 2018 1 repository listedHence, an object-to-motion convolutional neural network (OM-CNN) is developed to predict the intra-frame saliency for DeepVS, which is composed of the objectness and motion subnets.
-
23 Jan 2018 1 repository listedExisting video saliency datasets lack variety and generality of common dynamic scenes and fall short in covering challenging situations in unconstrained environments.
Syntology lines on 4 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections