{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/online-action-detection","title":"Online Action Detection","arxiv_id":"1604.06506","date":"2016-04-21","proceeding":null,"authors":["Roeland De Geest","Efstratios Gavves","Amir Ghodrati","Zhenyang Li","Cees Snoek","Tinne Tuytelaars"],"abstract":"In online action detection, the goal is to detect the start of an action in a\nvideo stream as soon as it happens. For instance, if a child is chasing a ball,\nan autonomous car should recognize what is going on and respond immediately.\nThis is a very challenging problem for four reasons. First, only partial\nactions are observed. Second, there is a large variability in negative data.\nThird, the start of the action is unknown, so it is unclear over what time\nwindow the information should be integrated. Finally, in real world data, large\nwithin-class variability exists. This problem has been addressed before, but\nonly to some extent. Our contributions to online action detection are\nthreefold. First, we introduce a realistic dataset composed of 27 episodes from\n6 popular TV series. The dataset spans over 16 hours of footage annotated with\n30 action classes, totaling 6,231 action instances. Second, we analyze and\ncompare various baseline methods, showing this is a challenging problem for\nwhich none of the methods provides a good solution. Third, we analyze the\nchange in performance when there is a variation in viewpoint, occlusion,\ntruncation, etc. We introduce an evaluation protocol for fair comparison. The\ndataset, the baselines and the models will all be made publicly available to\nencourage (much needed) further research on online action detection on\nrealistic data.","url_abs":"http://arxiv.org/abs/1604.06506v2","url_pdf":"http://arxiv.org/pdf/1604.06506v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"action-detection","task_name":"Action Detection"},{"task_slug":"online-action-detection","task_name":"Online Action Detection"}],"methods":[],"datasets_introduced":[{"slug":"tvseries","name":"TVSeries","full_name":""}],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1604.06506","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}