{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-learning-based-visual-saliency-prediction","title":"A Learning-Based Visual Saliency Prediction Model for Stereoscopic 3D Video (LBVS-3D)","arxiv_id":"1803.04842","date":"2018-03-13","proceeding":null,"authors":["Amin Banitalebi-Dehkordi","Mahsa T. Pourazad","Panos Nasiopoulos"],"abstract":"Over the past decade, many computational saliency prediction models have been\nproposed for 2D images and videos. Considering that the human visual system has\nevolved in a natural 3D environment, it is only natural to want to design\nvisual attention models for 3D content. Existing monocular saliency models are\nnot able to accurately predict the attentive regions when applied to 3D\nimage/video content, as they do not incorporate depth information. This paper\nexplores stereoscopic video saliency prediction by exploiting both low-level\nattributes such as brightness, color, texture, orientation, motion, and depth,\nas well as high-level cues such as face, person, vehicle, animal, text, and\nhorizon. Our model starts with a rough segmentation and quantifies several\nintuitive observations such as the effects of visual discomfort level, depth\nabruptness, motion acceleration, elements of surprise, size and compactness of\nthe salient regions, and emphasizing only a few salient objects in a scene. A\nnew fovea-based model of spatial distance between the image regions is adopted\nfor considering local and global feature calculations. To efficiently fuse the\nconspicuity maps generated by our method to one single saliency map that is\nhighly correlated with the eye-fixation data, a random forest based algorithm\nis utilized. The performance of the proposed saliency model is evaluated\nagainst the results of an eye-tracking experiment, which involved 24 subjects\nand an in-house database of 61 captured stereoscopic videos. Our stereo video\ndatabase as well as the eye-tracking data are publicly available along with\nthis paper. Experiment results show that the proposed saliency prediction\nmethod achieves competitive performance compared to the state-of-the-art\napproaches.","url_abs":"http://arxiv.org/abs/1803.04842v1","url_pdf":"http://arxiv.org/pdf/1803.04842v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"a-learning-based-visual-saliency-prediction","repo_url":"https://gitlab.com/abanitalebi/lbvs-3d","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"saliency-prediction","task_name":"Saliency Prediction"},{"task_slug":"video-saliency-prediction","task_name":"Video Saliency Prediction"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}