{"url":"/task/video-salient-object-detection","name":"Video Salient Object Detection","slug":"video-salient-object-detection","description_markdown":"Video salient object detection (VSOD) is significantly essential for understanding the underlying mechanism behind HVS during free-viewing in general and instrumental to a wide range of real-world applications, e.g., video segmentation, video captioning, video compression, autonomous driving, robotic interaction, weakly supervised attention. Besides its academic value and practical significance, VSOD presents great difficulties due to the challenges carried by video data (diverse motion patterns, occlusions, blur, large object deformations, etc.) and the inherent complexity of human visual attention behavior (i.e., selective attention allocation, attention shift) during dynamic scenes. Online benchmark: http://dpfan.net/davsod.\r\n\r\n<span style=\"color:grey; opacity: 0.6\">( Image credit: [Shifting More Attention to Video Salient Object Detection, CVPR2019-Best Paper Finalist](https://openaccess.thecvf.com/content_CVPR_2019/papers/Fan_Shifting_More_Attention_to_Video_Salient_Object_Detection_CVPR_2019_paper.pdf) )</span>","categories":[{"name":"Computer Vision","url":"/area/computer-vision"}],"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","slug_source":"archive_url"},"counts":{"papers_tagged":48,"papers_with_code":20,"benchmarks":10,"benchmark_tables_in_archive":10,"benchmark_tables_shown":10,"benchmark_tables_withheld_as_spam":0,"benchmark_definition":"a leaderboard table with at least one row; benchmark_tables_shown also counts the zero-row tables; benchmark_tables_in_archive adds the tables withheld as spam","datasets":4,"subtasks":0,"parent_tasks":2},"benchmarks":[{"leaderboard":"/sota/video-salient-object-detection-on-fbms-59","slug":"video-salient-object-detection-on-fbms-59","dataset":"FBMS-59","dataset_url":"/dataset/fbms-59","rows_in_archive":16,"metrics":["S-Measure","MAX E-MEASURE","MAX F-MEASURE","AVERAGE MAE"],"first_row_in_archive_order":{"model":"RealFlow","paper_title":"Transforming Static Images Using Generative Models for Video Salient Object Detection","paper_url":"/paper/transforming-static-images-using-generative","paper_date":"2024-11-21","arxiv_id":"2411.13975","code_links":[],"syntology":null}},{"leaderboard":"/sota/video-salient-object-detection-on-davis-2016","slug":"video-salient-object-detection-on-davis-2016","dataset":"DAVIS-2016","dataset_url":null,"rows_in_archive":11,"metrics":["S-Measure","MAX E-MEASURE","MAX F-MEASURE","AVERAGE MAE"],"first_row_in_archive_order":{"model":"RealFlow","paper_title":"Transforming Static Images Using Generative Models for Video Salient Object Detection","paper_url":"/paper/transforming-static-images-using-generative","paper_date":"2024-11-21","arxiv_id":"2411.13975","code_links":[],"syntology":null}},{"leaderboard":"/sota/video-salient-object-detection-on-visal","slug":"video-salient-object-detection-on-visal","dataset":"ViSal","dataset_url":"/dataset/visal","rows_in_archive":10,"metrics":["S-Measure","max E-measure","Average MAE"],"first_row_in_archive_order":{"model":"RealFlow","paper_title":"Transforming Static Images Using Generative Models for Video Salient Object Detection","paper_url":"/paper/transforming-static-images-using-generative","paper_date":"2024-11-21","arxiv_id":"2411.13975","code_links":[],"syntology":null}},{"leaderboard":"/sota/video-salient-object-detection-on-davsod","slug":"video-salient-object-detection-on-davsod","dataset":"DAVSOD-easy35","dataset_url":null,"rows_in_archive":9,"metrics":["S-Measure","max F-Measure","Average MAE","max E-Measure"],"first_row_in_archive_order":{"model":"RealFlow","paper_title":"Transforming Static Images Using Generative Models for Video Salient Object Detection","paper_url":"/paper/transforming-static-images-using-generative","paper_date":"2024-11-21","arxiv_id":"2411.13975","code_links":[],"syntology":null}},{"leaderboard":"/sota/video-salient-object-detection-on-vos-t","slug":"video-salient-object-detection-on-vos-t","dataset":"VOS-T","dataset_url":null,"rows_in_archive":9,"metrics":["S-Measure","max E-measure","Average MAE"],"first_row_in_archive_order":{"model":"RCRNet+NER","paper_title":"Semi-Supervised Video Salient Object Detection Using Pseudo-Labels","paper_url":"/paper/semi-supervised-video-salient-object","paper_date":"2019-08-12","arxiv_id":"1908.04051","code_links":[{"title":"Kinpzz/RCRNet-Pytorch","url":"https://github.com/Kinpzz/RCRNet-Pytorch"}],"syntology":null}},{"leaderboard":"/sota/video-salient-object-detection-on-davsod-1","slug":"video-salient-object-detection-on-davsod-1","dataset":"DAVSOD-Normal25","dataset_url":null,"rows_in_archive":8,"metrics":["S-Measure","max E-measure","Average MAE"],"first_row_in_archive_order":{"model":"SSAV","paper_title":"Shifting More Attention to Video Salient Object Detection","paper_url":"/paper/shifting-more-attention-to-video-salient","paper_date":"2019-06-01","arxiv_id":null,"code_links":[{"title":"DengPingFan/DAVSOD","url":"https://github.com/DengPingFan/DAVSOD"}],"syntology":null}},{"leaderboard":"/sota/video-salient-object-detection-on-davsod-2","slug":"video-salient-object-detection-on-davsod-2","dataset":"DAVSOD-Difficult20","dataset_url":null,"rows_in_archive":8,"metrics":["S-Measure","max E-measure","Average MAE"],"first_row_in_archive_order":{"model":"SSAV","paper_title":"Shifting More Attention to Video Salient Object Detection","paper_url":"/paper/shifting-more-attention-to-video-salient","paper_date":"2019-06-01","arxiv_id":null,"code_links":[{"title":"DengPingFan/DAVSOD","url":"https://github.com/DengPingFan/DAVSOD"}],"syntology":null}},{"leaderboard":"/sota/video-salient-object-detection-on-mcl","slug":"video-salient-object-detection-on-mcl","dataset":"MCL","dataset_url":null,"rows_in_archive":8,"metrics":["S-Measure","MAX E-MEASURE","MAX F-MEASURE","AVERAGE MAE"],"first_row_in_archive_order":{"model":"PDB","paper_title":"Pyramid Dilated Deeper ConvLSTM for Video Salient Object Detection","paper_url":"/paper/pyramid-dilated-deeper-convlstm-for-video","paper_date":"2018-09-01","arxiv_id":null,"code_links":[],"syntology":null}},{"leaderboard":"/sota/video-salient-object-detection-on-segtrack-v2","slug":"video-salient-object-detection-on-segtrack-v2","dataset":"SegTrack v2","dataset_url":"/dataset/segtrack-v2-1","rows_in_archive":8,"metrics":["S-Measure","max E-measure","MAX F-MEASURE","AVERAGE MAE"],"first_row_in_archive_order":{"model":"UFO","paper_title":"A Unified Transformer Framework for Group-based Segmentation: Co-Segmentation, Co-Saliency Detection and Video Salient Object Detection","paper_url":"/paper/a-unified-transformer-framework-for-group","paper_date":"2022-03-09","arxiv_id":"2203.04708","code_links":[{"title":"suyukun666/UFO","url":"https://github.com/suyukun666/UFO"}],"syntology":{"n":13,"n_ran":0,"n_unverified":13,"n_pointer_only":0}}},{"leaderboard":"/sota/video-salient-object-detection-on-uvsd","slug":"video-salient-object-detection-on-uvsd","dataset":"UVSD","dataset_url":null,"rows_in_archive":8,"metrics":["S-Measure","max E-measure","Average MAE"],"first_row_in_archive_order":{"model":"PDB","paper_title":"Pyramid Dilated Deeper ConvLSTM for Video Salient Object Detection","paper_url":"/paper/pyramid-dilated-deeper-convlstm-for-video","paper_date":"2018-09-01","arxiv_id":null,"code_links":[],"syntology":null}}],"datasets":[{"url":"/dataset/fbms","name":"FBMS","full_name":"Freiburg-Berkeley Motion Segmentation","num_papers_in_archive":126},{"url":"/dataset/segtrack-v2-1","name":"SegTrack-v2","full_name":"","num_papers_in_archive":107},{"url":"/dataset/fbms-59","name":"FBMS-59","full_name":"Freiburg-Berkeley Motion Segmentation","num_papers_in_archive":19},{"url":"/dataset/visal","name":"ViSal","full_name":"","num_papers_in_archive":10}],"subtasks":[],"parent_tasks":[{"url":"/task/salient-object-detection","name":"RGB Salient Object Detection"},{"url":"/task/video-object-segmentation","name":"Video Object Segmentation"}],"papers":{"order":"repositories listed in the archive (desc), then date (desc); the archive holds no stars","population":"papers tagged with this task that list at least one repository in the archive","shown":20,"of":20,"tagged_in_all":48,"items":[{"url":"/paper/depth-cooperated-trimodal-network-for-video","title":"Depth-Cooperated Trimodal Network for Video Salient Object Detection","date":"2022-02-12","arxiv_id":"2202.06060","repositories_listed":2,"syntology":null},{"url":"/paper/motion-guided-attention-for-video-salient","title":"Motion Guided Attention for Video Salient Object Detection","date":"2019-09-16","arxiv_id":"1909.07061","repositories_listed":2,"syntology":null},{"url":"/paper/learning-motion-and-temporal-cues-for","title":"Learning Motion and Temporal Cues for Unsupervised Video Object Segmentation","date":"2025-01-14","arxiv_id":"2501.07806","repositories_listed":1,"syntology":null},{"url":"/paper/vidsod-100-a-new-dataset-and-a-baseline-model","title":"ViDSOD-100: A New Dataset and a Baseline Model for RGB-D Video Salient Object Detection","date":"2024-06-18","arxiv_id":"2406.12536","repositories_listed":1,"syntology":null},{"url":"/paper/depth-quality-inspired-feature-manipulation-1","title":"Depth Quality-Inspired Feature Manipulation for Efficient RGB-D and Video Salient Object Detection","date":"2022-08-08","arxiv_id":"2208.03918","repositories_listed":1,"syntology":null},{"url":"/paper/motion-aware-memory-network-for-fast-video","title":"Motion-aware Memory Network for Fast Video Salient Object Detection","date":"2022-08-01","arxiv_id":"2208.00946","repositories_listed":1,"syntology":null},{"url":"/paper/hierarchical-feature-alignment-network-for","title":"Hierarchical Feature Alignment Network for Unsupervised Video Object Segmentation","date":"2022-07-18","arxiv_id":"2207.08485","repositories_listed":1,"syntology":null},{"url":"/paper/learning-video-salient-object-detection","title":"Learning Video Salient Object Detection Progressively from Unlabeled Videos","date":"2022-04-05","arxiv_id":"2204.02008","repositories_listed":1,"syntology":null},{"url":"/paper/a-unified-transformer-framework-for-group","title":"A Unified Transformer Framework for Group-based Segmentation: Co-Segmentation, Co-Saliency Detection and Video Salient Object Detection","date":"2022-03-09","arxiv_id":"2203.04708","repositories_listed":1,"syntology":{"n":13,"n_ran":0,"n_unverified":13,"n_pointer_only":0}},{"url":"/paper/full-duplex-strategy-for-video-object","title":"Full-Duplex Strategy for Video Object Segmentation","date":"2021-08-06","arxiv_id":"2108.03151","repositories_listed":1,"syntology":{"n":8,"n_ran":7,"n_unverified":1,"n_pointer_only":0}},{"url":"/paper/video-salient-object-detection-via-adaptive","title":"Video Salient Object Detection via Adaptive Local-Global Refinement","date":"2021-04-29","arxiv_id":"2104.14360","repositories_listed":1,"syntology":null},{"url":"/paper/weakly-supervised-video-salient-object","title":"Weakly Supervised Video Salient Object Detection","date":"2021-04-06","arxiv_id":"2104.02391","repositories_listed":1,"syntology":null},{"url":"/paper/dynamic-context-sensitive-filtering-network","title":"Dynamic Context-Sensitive Filtering Network for Video Salient Object Detection","date":"2021-01-01","arxiv_id":null,"repositories_listed":1,"syntology":null},{"url":"/paper/ds-net-dynamic-spatiotemporal-network-for","title":"DS-Net: Dynamic Spatiotemporal Network for Video Salient Object Detection","date":"2020-12-09","arxiv_id":"2012.04886","repositories_listed":1,"syntology":null},{"url":"/paper/exploring-rich-and-efficient-spatial-temporal","title":"Exploring Rich and Efficient Spatial Temporal Interactions for Real Time Video Salient Object Detection","date":"2020-08-07","arxiv_id":"2008.02973","repositories_listed":1,"syntology":null},{"url":"/paper/a-novel-video-salient-object-detection-method","title":"A Novel Video Salient Object Detection Method via Semi-supervised Motion Quality Perception","date":"2020-08-07","arxiv_id":"2008.02966","repositories_listed":1,"syntology":null},{"url":"/paper/semi-supervised-video-salient-object","title":"Semi-Supervised Video Salient Object Detection Using Pseudo-Labels","date":"2019-08-12","arxiv_id":"1908.04051","repositories_listed":1,"syntology":null},{"url":"/paper/shifting-more-attention-to-video-salient","title":"Shifting More Attention to Video Salient Object Detection","date":"2019-06-01","arxiv_id":null,"repositories_listed":1,"syntology":null},{"url":"/paper/structure-measure-a-new-way-to-evaluate","title":"Structure-measure: A New Way to Evaluate Foreground Maps","date":"2017-08-02","arxiv_id":"1708.00786","repositories_listed":1,"syntology":null},{"url":"/paper/real-time-salient-object-detection-with-a","title":"Real-Time Salient Object Detection With a Minimum Spanning Tree","date":"2016-06-01","arxiv_id":null,"repositories_listed":1,"syntology":null}],"syntology_records":2,"syntology_note":"a paper without a record is not a recorded non-run: it may lack an arXiv id or simply be absent from the graph layer"},"description_links":{"kept":0,"unwrapped_to_text":0,"bare_urls_linked":0,"relative_images_dropped":0,"rule":"internal links are kept only when the target slug exists in the catalog"},"syntology":{"read_at":"2026-09-24T18:15:14+00:00","claim":"Per-sample execution status on synthesized fixtures ('ran N of M samples'); not a correctness claim and not a ranking signal.","status_vocabulary":{"ran_honours":"ran, honoured the contract we drafted","ran_violates":"ran, violated the contract we drafted","ran_draft_wrong":"ran; our contract draft was wrong, not the code","ran_fixture":"ran; our fixture could not drive it","ran":"ran on a synthesized input","unverified":"unverified (harvested, no recorded run)"}},"not_shown":{"libraries":"the archive has no per-task library table","trend_sparklines":"the Trend column of the benchmarks table was a rendered image; it is not in the archive","social_and_latest_sorts":"stars and social signals are not in the archive"}}