{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/pooled-motion-features-for-first-person","title":"Pooled Motion Features for First-Person Videos","arxiv_id":"1412.6505","date":"2014-12-19","proceeding":"CVPR 2015 6","authors":["M. S. Ryoo","Brandon Rothrock","Larry Matthies"],"abstract":"In this paper, we present a new feature representation for first-person\nvideos. In first-person video understanding (e.g., activity recognition), it is\nvery important to capture both entire scene dynamics (i.e., egomotion) and\nsalient local motion observed in videos. We describe a representation framework\nbased on time series pooling, which is designed to abstract\nshort-term/long-term changes in feature descriptor elements. The idea is to\nkeep track of how descriptor values are changing over time and summarize them\nto represent motion in the activity video. The framework is general, handling\nany types of per-frame feature descriptors including conventional motion\ndescriptors like histogram of optical flows (HOF) as well as appearance\ndescriptors from more recent convolutional neural networks (CNN). We\nexperimentally confirm that our approach clearly outperforms previous feature\nrepresentations including bag-of-visual-words and improved Fisher vector (IFV)\nwhen using identical underlying feature descriptors. We also confirm that our\nfeature representation has superior performance to existing state-of-the-art\nfeatures like local spatio-temporal features and Improved Trajectory Features\n(originally developed for 3rd-person videos) when handling first-person videos.\nMultiple first-person activity datasets were tested under various settings to\nconfirm these findings.","url_abs":"http://arxiv.org/abs/1412.6505v2","url_pdf":"http://arxiv.org/pdf/1412.6505v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"pooled-motion-features-for-first-person","repo_url":"https://github.com/mryoo/pooled_time_series","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"activity-recognition","task_name":"Activity Recognition"},{"task_slug":"activity-recognition-in-videos","task_name":"Activity Recognition In Videos"},{"task_slug":"time-series-1","task_name":"Time Series"},{"task_slug":"time-series","task_name":"Time Series Analysis"},{"task_slug":"video-understanding","task_name":"Video Understanding"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}