{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/egocentric-video-description-based-on","title":"Egocentric Video Description based on Temporally-Linked Sequences","arxiv_id":"1704.02163","date":"2017-04-07","proceeding":null,"authors":["Marc Bolaños","Álvaro Peris","Francisco Casacuberta","Sergi Soler","Petia Radeva"],"abstract":"Egocentric vision consists in acquiring images along the day from a first\nperson point-of-view using wearable cameras. The automatic analysis of this\ninformation allows to discover daily patterns for improving the quality of life\nof the user. A natural topic that arises in egocentric vision is storytelling,\nthat is, how to understand and tell the story relying behind the pictures. In\nthis paper, we tackle storytelling as an egocentric sequences description\nproblem. We propose a novel methodology that exploits information from\ntemporally neighboring events, matching precisely the nature of egocentric\nsequences. Furthermore, we present a new method for multimodal data fusion\nconsisting on a multi-input attention recurrent network. We also publish the\nfirst dataset for egocentric image sequences description, consisting of 1,339\nevents with 3,991 descriptions, from 55 days acquired by 11 people.\nFurthermore, we prove that our proposal outperforms classical attentional\nencoder-decoder methods for video description.","url_abs":"http://arxiv.org/abs/1704.02163v3","url_pdf":"http://arxiv.org/pdf/1704.02163v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"egocentric-video-description-based-on","repo_url":"https://github.com/MarcBS/TMA","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"decoder","task_name":"Decoder"},{"task_slug":"video-description","task_name":"Video Description"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}