Papers › Transforming Static Images Using Generative Models for Video Salient Object Detection

Transforming Static Images Using Generative Models for Video Salient Object Detection

21 Nov 2024arXiv:2411.13975archive 2025-07-28

Suhwan Cho, Minhyeok Lee, Jungho Lee, Sangyoun Lee

In many video processing tasks, leveraging large-scale image datasets is a common strategy, as image data is more abundant and facilitates comprehensive knowledge transfer. A typical approach for simulating video from static images involves applying spatial transformations, such as affine transformations and spline warping, to create sequences that mimic temporal progression. However, in tasks like video salient object detection, where both appearance and motion cues are critical, these basic image-to-video techniques fail to produce realistic optical flows that capture the independent motion properties of each object. In this study, we show that image-to-video diffusion models can generate realistic transformations of static images while understanding the contextual relationships between image components. This ability allows the model to generate plausible optical flows, preserving semantic integrity while reflecting the independent motion of scene elements. By augmenting individual images in this way, we create large-scale image-flow pairs that significantly enhance model training. Our approach achieves state-of-the-art performance across all public benchmark datasets, outperforming existing approaches.

PaperPDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Object DetectionSalient Object DetectionTransfer LearningVideo Salient Object Detectionobject-detection

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Video Salient Object Detection DAVIS-2016 RealFlow AVERAGE MAE 0.010 #1 of 11 Archive leaderboard report
Video Salient Object Detection DAVIS-2016 RealFlow MAX F-MEASURE 0.939 #1 of 11 Archive leaderboard report
Video Salient Object Detection DAVIS-2016 RealFlow S-Measure 0.945 #1 of 11 Archive leaderboard report
Video Salient Object Detection DAVSOD-easy35 RealFlow Average MAE 0.066 #1 of 9 Archive leaderboard report
Video Salient Object Detection DAVSOD-easy35 RealFlow S-Measure 0.803 #1 of 9 Archive leaderboard report
Video Salient Object Detection DAVSOD-easy35 RealFlow max F-Measure 0.732 #1 of 9 Archive leaderboard report
Video Salient Object Detection FBMS-59 RealFlow AVERAGE MAE 0.028 #1 of 16 Archive leaderboard report
Video Salient Object Detection FBMS-59 RealFlow MAX F-MEASURE 0.906 #1 of 16 Archive leaderboard report
Video Salient Object Detection FBMS-59 RealFlow S-Measure 0.926 #1 of 16 Archive leaderboard report
Video Salient Object Detection ViSal RealFlow Average MAE 0.010 #1 of 10 Archive leaderboard report
Video Salient Object Detection ViSal RealFlow S-Measure 0.962 #1 of 10 Archive leaderboard report
Video Salient Object Detection ViSal RealFlow max E-measure 0.966 #1 of 10 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Diffusion

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections