{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-visual-storylines-with-skipping","title":"Learning Visual Storylines with Skipping Recurrent Neural Networks","arxiv_id":"1604.04279","date":"2016-04-14","proceeding":null,"authors":["Gunnar A. Sigurdsson","Xinlei Chen","Abhinav Gupta"],"abstract":"What does a typical visit to Paris look like? Do people first take photos of\nthe Louvre and then the Eiffel Tower? Can we visually model a temporal event\nlike \"Paris Vacation\" using current frameworks? In this paper, we explore how\nwe can automatically learn the temporal aspects, or storylines of visual\nconcepts from web data. Previous attempts focus on consecutive image-to-image\ntransitions and are unsuccessful at recovering the long-term underlying story.\nOur novel Skipping Recurrent Neural Network (S-RNN) model does not attempt to\npredict each and every data point in the sequence, like classic RNNs. Rather,\nS-RNN uses a framework that skips through the images in the photo stream to\nexplore the space of all ordered subsets of the albums via an efficient\nsampling procedure. This approach reduces the negative impact of strong\nshort-term correlations, and recovers the latent story more accurately. We show\nhow our learned storylines can be used to analyze, predict, and summarize photo\nalbums from Flickr. Our experimental results provide strong qualitative and\nquantitative evidence that S-RNN is significantly better than other candidate\nmethods such as LSTMs on learning long-term correlations and recovering latent\nstorylines. Moreover, we show how storylines can help machines better\nunderstand and summarize photo streams by inferring a brief personalized story\nof each individual album.","url_abs":"http://arxiv.org/abs/1604.04279v2","url_pdf":"http://arxiv.org/pdf/1604.04279v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"learning-visual-storylines-with-skipping","repo_url":"https://github.com/gsig/srnn","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}