{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/predicting-video-saliency-with-object-to","title":"Predicting Video Saliency with Object-to-Motion CNN and Two-layer Convolutional LSTM","arxiv_id":"1709.06316","date":"2017-09-19","proceeding":null,"authors":["Lai Jiang","Mai Xu","Zulin Wang"],"abstract":"Over the past few years, deep neural networks (DNNs) have exhibited great\nsuccess in predicting the saliency of images. However, there are few works that\napply DNNs to predict the saliency of generic videos. In this paper, we propose\na novel DNN-based video saliency prediction method. Specifically, we establish\na large-scale eye-tracking database of videos (LEDOV), which provides\nsufficient data to train the DNN models for predicting video saliency. Through\nthe statistical analysis of our LEDOV database, we find that human attention is\nnormally attracted by objects, particularly moving objects or the moving parts\nof objects. Accordingly, we propose an object-to-motion convolutional neural\nnetwork (OM-CNN) to learn spatio-temporal features for predicting the\nintra-frame saliency via exploring the information of both objectness and\nobject motion. We further find from our database that there exists a temporal\ncorrelation of human attention with a smooth saliency transition across video\nframes. Therefore, we develop a two-layer convolutional long short-term memory\n(2C-LSTM) network in our DNN-based method, using the extracted features of\nOM-CNN as the input. Consequently, the inter-frame saliency maps of videos can\nbe generated, which consider the transition of attention across video frames.\nFinally, the experimental results show that our method advances the\nstate-of-the-art in video saliency prediction.","url_abs":"http://arxiv.org/abs/1709.06316v3","url_pdf":"http://arxiv.org/pdf/1709.06316v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"predicting-video-saliency-with-object-to","repo_url":"https://github.com/remega/LEDOV-eye-tracking-database","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"saliency-prediction","task_name":"Saliency Prediction"},{"task_slug":"video-saliency-prediction","task_name":"Video Saliency Prediction"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1709.06316","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}