{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/viena2-a-driving-anticipation-dataset","title":"VIENA2: A Driving Anticipation Dataset","arxiv_id":"1810.09044","date":"2018-10-22","proceeding":null,"authors":["Mohammad Sadegh Aliakbarian","Fatemeh Sadat Saleh","Mathieu Salzmann","Basura Fernando","Lars Petersson","Lars Andersson"],"abstract":"Action anticipation is critical in scenarios where one needs to react before\nthe action is finalized. This is, for instance, the case in automated driving,\nwhere a car needs to, e.g., avoid hitting pedestrians and respect traffic\nlights. While solutions have been proposed to tackle subsets of the driving\nanticipation tasks, by making use of diverse, task-specific sensors, there is\nno single dataset or framework that addresses them all in a consistent manner.\nIn this paper, we therefore introduce a new, large-scale dataset, called\nVIENA2, covering 5 generic driving scenarios, with a total of 25 distinct\naction classes. It contains more than 15K full HD, 5s long videos acquired in\nvarious driving conditions, weathers, daytimes and environments, complemented\nwith a common and realistic set of sensor measurements. This amounts to more\nthan 2.25M frames, each annotated with an action label, corresponding to 600\nsamples per action class. We discuss our data acquisition strategy and the\nstatistics of our dataset, and benchmark state-of-the-art action anticipation\ntechniques, including a new multi-modal LSTM architecture with an effective\nloss function for action anticipation in driving scenarios.","url_abs":"http://arxiv.org/abs/1810.09044v2","url_pdf":"http://arxiv.org/pdf/1810.09044v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"action-anticipation","task_name":"Action Anticipation"}],"methods":[{"method_slug":"lstm","method_name":"LSTM"},{"method_slug":"sigmoid-activation","method_name":"Sigmoid Activation"},{"method_slug":"tanh-activation","method_name":"Tanh Activation"}],"datasets_introduced":[{"slug":"viena2","name":"VIENA2","full_name":""}],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1810.09044","atlas_url":"https://app.syntology.ai/?focus=1810.09044","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}