{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/shifting-more-attention-to-video-salient","title":"Shifting More Attention to Video Salient Object Detection","arxiv_id":null,"date":"2019-06-01","proceeding":"CVPR 2019 6","authors":["Deng-Ping Fan"," Wenguan Wang"," Ming-Ming Cheng"," Jianbing Shen"],"abstract":"The last decade has witnessed a growing interest in video salient object detection (VSOD). However, the research community long-term lacked a well-established VSOD dataset representative of real dynamic scenes with high-quality annotations. To address this issue, we elaborately collected a visual-attention-consistent Densely Annotated VSOD (DAVSOD) dataset, which contains 226 videos with 23,938 frames that cover diverse realistic-scenes, objects, instances and motions. With corresponding real human eye-fixation data, we obtain precise ground-truths. This is the first work that explicitly emphasizes the challenge of saliency shift, i.e., the video salient object(s) may dynamically change. To further contribute the community a complete benchmark, we systematically assess 17 representative VSOD algorithms over seven existing VSOD datasets and our DAVSOD with totally  84K frames (largest-scale). Utilizing three famous metrics, we then present a comprehensive and insightful performance analysis. Furthermore, we propose a baseline model. It is equipped with a saliency shift- aware convLSTM, which can efficiently capture video saliency dynamics through learning human attention-shift behavior. Extensive experiments open up promising future directions for model development and comparison.\r","url_abs":"http://openaccess.thecvf.com/content_CVPR_2019/html/Fan_Shifting_More_Attention_to_Video_Salient_Object_Detection_CVPR_2019_paper.html","url_pdf":"http://openaccess.thecvf.com/content_CVPR_2019/papers/Fan_Shifting_More_Attention_to_Video_Salient_Object_Detection_CVPR_2019_paper.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"shifting-more-attention-to-video-salient","repo_url":"https://github.com/DengPingFan/DAVSOD","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"tf","reach":null}],"tasks":[{"task_slug":"object","task_name":"Object"},{"task_slug":"object-detection","task_name":"Object Detection"},{"task_slug":"salient-object-detection","task_name":"RGB Salient Object Detection"},{"task_slug":"salient-object-detection-1","task_name":"Salient Object Detection"},{"task_slug":"video-object-segmentation","task_name":"Video Object Segmentation"},{"task_slug":"video-salient-object-detection","task_name":"Video Salient Object Detection"},{"task_slug":"object-detection-1","task_name":"object-detection"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/video-salient-object-detection-on-davis-2016","task":"Video Salient Object Detection","dataset":"DAVIS-2016","model":"SSAV","rank_in_archive_order":3,"of":11,"metrics":{"AVERAGE MAE":"0.028","MAX E-MEASURE":"0.948","MAX F-MEASURE":"0.861","S-Measure":"0.893"},"uses_additional_data":true},{"leaderboard":"/sota/video-salient-object-detection-on-davsod-2","task":"Video Salient Object Detection","dataset":"DAVSOD-Difficult20","model":"SSAV","rank_in_archive_order":1,"of":8,"metrics":{"Average MAE":"0.114","S-Measure":"0.619","max E-measure":"0.696"},"uses_additional_data":false},{"leaderboard":"/sota/video-salient-object-detection-on-davsod-1","task":"Video Salient Object Detection","dataset":"DAVSOD-Normal25","model":"SSAV","rank_in_archive_order":1,"of":8,"metrics":{"Average MAE":"0.117","S-Measure":"0.661","max E-measure":"0.723"},"uses_additional_data":false},{"leaderboard":"/sota/video-salient-object-detection-on-davsod","task":"Video Salient Object Detection","dataset":"DAVSOD-easy35","model":"SSAV","rank_in_archive_order":2,"of":9,"metrics":{"Average MAE":"0.084","S-Measure":"0.755","max E-Measure":"0.806","max F-Measure":"0.659"},"uses_additional_data":false},{"leaderboard":"/sota/video-salient-object-detection-on-fbms-59","task":"Video Salient Object Detection","dataset":"FBMS-59","model":"SSAV","rank_in_archive_order":3,"of":16,"metrics":{"AVERAGE MAE":"0.040","MAX E-MEASURE":"0.926","MAX F-MEASURE":"0.865","S-Measure":"0.879"},"uses_additional_data":true},{"leaderboard":"/sota/video-salient-object-detection-on-mcl","task":"Video Salient Object Detection","dataset":"MCL","model":"SSAV","rank_in_archive_order":2,"of":8,"metrics":{"AVERAGE MAE":"0.026","MAX E-MEASURE":"0.889","MAX F-MEASURE":"0.773","S-Measure":"0.819"},"uses_additional_data":true},{"leaderboard":"/sota/video-salient-object-detection-on-segtrack-v2","task":"Video Salient Object Detection","dataset":"SegTrack v2","model":"SSAV","rank_in_archive_order":3,"of":8,"metrics":{"AVERAGE MAE":"0.023","MAX F-MEASURE":"0.801","S-Measure":"0.850","max E-measure":"0.917"},"uses_additional_data":false},{"leaderboard":"/sota/video-salient-object-detection-on-uvsd","task":"Video Salient Object Detection","dataset":"UVSD","model":"SSAV","rank_in_archive_order":2,"of":8,"metrics":{"Average MAE":"0.025","S-Measure":"0.860","max E-measure":"0.939"},"uses_additional_data":true},{"leaderboard":"/sota/video-salient-object-detection-on-vos-t","task":"Video Salient Object Detection","dataset":"VOS-T","model":"SSAV","rank_in_archive_order":2,"of":9,"metrics":{"Average MAE":"0.074","S-Measure":"0.819","max E-measure":"0.839"},"uses_additional_data":false},{"leaderboard":"/sota/video-salient-object-detection-on-visal","task":"Video Salient Object Detection","dataset":"ViSal","model":"SSAV","rank_in_archive_order":3,"of":10,"metrics":{"Average MAE":"0.021","S-Measure":"0.942","max E-measure":"0.980"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}