{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/visual-saliency-based-on-scale-space-analysis","title":"Visual Saliency Based on Scale-Space Analysis in the Frequency Domain","arxiv_id":"1605.01999","date":"2016-05-06","proceeding":null,"authors":["Jian Li","Martin Levine","Xiangjing An","Xin Xu","Hangen He"],"abstract":"We address the issue of visual saliency from three perspectives. First, we\nconsider saliency detection as a frequency domain analysis problem. Second, we\nachieve this by employing the concept of {\\it non-saliency}. Third, we\nsimultaneously consider the detection of salient regions of different size. The\npaper proposes a new bottom-up paradigm for detecting visual saliency,\ncharacterized by a scale-space analysis of the amplitude spectrum of natural\nimages. We show that the convolution of the {\\it image amplitude spectrum} with\na low-pass Gaussian kernel of an appropriate scale is equivalent to such an\nimage saliency detector. The saliency map is obtained by reconstructing the 2-D\nsignal using the original phase and the amplitude spectrum, filtered at a scale\nselected by minimizing saliency map entropy. A Hypercomplex Fourier Transform\nperforms the analysis in the frequency domain. Using available databases, we\ndemonstrate experimentally that the proposed model can predict human fixation\ndata. We also introduce a new image database and use it to show that the\nsaliency detector can highlight both small and large salient regions, as well\nas inhibit repeated distractors in cluttered images. In addition, we show that\nit is able to predict salient regions on which people focus their attention.","url_abs":"http://arxiv.org/abs/1605.01999v1","url_pdf":"http://arxiv.org/pdf/1605.01999v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"saliency-detection","task_name":"Saliency Detection"},{"task_slug":"video-saliency-detection","task_name":"Video Saliency Detection"}],"methods":[{"method_slug":"convolution","method_name":"Convolution"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/video-saliency-detection-on-msu-video","task":"Video Saliency Detection","dataset":"MSU Video Saliency Prediction","model":"HFT","rank_in_archive_order":11,"of":14,"metrics":{"AUC-J":"0.814","CC":"0.577","FPS":"3.63","KLDiv":"0.698","NSS":"1.35","SIM":"0.550"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}