Browse State-of-the-Art › Video Salient Object Detection
Video Salient Object Detection
20 papers with code · 10 benchmarks · 4 datasets archive 2025-07-28
Video salient object detection (VSOD) is significantly essential for understanding the underlying mechanism behind HVS during free-viewing in general and instrumental to a wide range of real-world applications, e.g., video segmentation, video captioning, video compression, autonomous driving, robotic interaction, weakly supervised attention. Besides its academic value and practical significance, VSOD presents great difficulties due to the challenges carried by video data (diverse motion patterns, occlusions, blur, large object deformations, etc.) and the inherent complexity of human visual attention behavior (i.e., selective attention allocation, attention shift) during dynamic scenes. Online benchmark: http://dpfan.net/davsod.
( Image credit: Shifting More Attention to Video Salient Object Detection, CVPR2019-Best Paper Finalist )
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
10 leaderboard tables shown for this task, 10 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
4 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
20 shown of 20 papers with code (48 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
12 Feb 2022 2 repositories listedHowever, existing video salient object detection (VSOD) methods only utilize spatiotemporal information and seldom exploit depth information for detection.
-
16 Sep 2019 2 repositories listedIn this paper, we develop a multi-task motion guided video salient object detection network, which learns to accomplish two sub-tasks using two sub-networks, one sub-network for salient object detection in still images…
-
14 Jan 2025 1 repository listedIn this paper, we address the challenges in unsupervised video object segmentation (UVOS) by proposing an efficient algorithm, termed MTNet, which concurrently exploits motion and temporal cues.
-
18 Jun 2024 1 repository listedHowever, the existing salient object detection (SOD) works only focus on either static RGB-D images or RGB videos, ignoring the collaborating of RGB-D and video information.
-
8 Aug 2022 1 repository listedInspired by the fact that depth quality is a key factor influencing the accuracy, we propose an efficient depth quality-inspired feature manipulation (DQFM) process, which can dynamically filter depth features according…
-
1 Aug 2022 1 repository listedMoreover, inspired by the boundary supervision commonly used in image salient object detection (ISOD), we design a motion-aware loss for predicting object boundary motion and simultaneously perform multitask learning…
-
18 Jul 2022 1 repository listedOptical flow is an easily conceived and precious cue for advancing unsupervised video object segmentation (UVOS).
-
5 Apr 2022 1 repository listedRecent deep learning-based video salient object detection (VSOD) has achieved some breakthrough, but these methods rely on expensive annotated videos with pixel-wise annotations, weak annotations, or part of the…
-
9 Mar 2022 1 repository listed Syntology ran 0 of 13 samples · 13 unverifiedBesides, they fail to take full advantage of the cues among inter- and intra-feature within a group of images.
-
6 Aug 2021 1 repository listed Syntology ran 7 of 8 samples · 1 unverifiedPrevious video object segmentation approaches mainly focus on using simplex solutions between appearance and motion, limiting feature collaboration efficiency among and across these two cues.
-
29 Apr 2021 1 repository listedDespite their simplicity, such fusion strategies may introduce feature redundancy, and also fail to fully exploit the relationship between multi-level features extracted from both spatial and temporal domains.
-
6 Apr 2021 1 repository listedSignificant performance improvement has been achieved for fully-supervised video salient object detection with the pixel-wise labeled training datasets, which are time-consuming and expensive to obtain.
-
1 Jan 2021 1 repository listedOur bidirectional dynamic fusion strategy encourages the interaction of spatial and temporal information in a dynamic manner.
-
9 Dec 2020 1 repository listedIn this paper, we investigate the complimentary roles of spatial and temporal information and propose a novel dynamic spatiotemporal network (DS-Net) for more effective fusion of spatiotemporal information.
-
7 Aug 2020 1 repository listedIn this way, even though the overall video saliency quality is heavily dependent on its spatial branch, however, the performance of the temporal branch still matter.
-
7 Aug 2020 1 repository listedConsequently, we can achieve a significant performance improvement by using this new training set to start a new round of network training.
-
12 Aug 2019 1 repository listedSpecifically, we present an effective video saliency detector that consists of a spatial refinement network and a spatiotemporal module.
-
1 Jun 2019 1 repository listedThis is the first work that explicitly emphasizes the challenge of saliency shift, i.
-
2 Aug 2017 1 repository listedOur new measure simultaneously evaluates region-aware and object-aware structural similarity between a SM and a GT map.
-
1 Jun 2016 1 repository listedIn this paper, we present a real-time salient object detection system based on the minimum spanning tree.
Syntology lines on 2 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections