{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/im2flow-motion-hallucination-from-static","title":"Im2Flow: Motion Hallucination from Static Images for Action Recognition","arxiv_id":"1712.04109","date":"2017-12-12","proceeding":"CVPR 2018 6","authors":["Ruohan Gao","Bo Xiong","Kristen Grauman"],"abstract":"Existing methods to recognize actions in static images take the images at\ntheir face value, learning the appearances---objects, scenes, and body\nposes---that distinguish each action class. However, such models are deprived\nof the rich dynamic structure and motions that also define human activity. We\npropose an approach that hallucinates the unobserved future motion implied by a\nsingle snapshot to help static-image action recognition. The key idea is to\nlearn a prior over short-term dynamics from thousands of unlabeled videos,\ninfer the anticipated optical flow on novel static images, and then train\ndiscriminative models that exploit both streams of information. Our main\ncontributions are twofold. First, we devise an encoder-decoder convolutional\nneural network and a novel optical flow encoding that can translate a static\nimage into an accurate flow map. Second, we show the power of hallucinated flow\nfor recognition, successfully transferring the learned motion into a standard\ntwo-stream network for activity recognition. On seven datasets, we demonstrate\nthe power of the approach. It not only achieves state-of-the-art accuracy for\ndense optical flow prediction, but also consistently enhances recognition of\nactions and dynamic scenes.","url_abs":"http://arxiv.org/abs/1712.04109v3","url_pdf":"http://arxiv.org/pdf/1712.04109v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"im2flow-motion-hallucination-from-static","repo_url":"https://github.com/YashNita/Separate-Object-Sounds-by-Watching-Unlabeled-Video-PyTorch-","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"im2flow-motion-hallucination-from-static","repo_url":"https://github.com/rhgao/Deep-MIML-Network","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"im2flow-motion-hallucination-from-static","repo_url":"https://github.com/rhgao/Im2Flow","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"torch","reach":null},{"paper_slug":"im2flow-motion-hallucination-from-static","repo_url":"https://github.com/rhgao/separating-object-sounds","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"activity-recognition","task_name":"Activity Recognition"},{"task_slug":"decoder","task_name":"Decoder"},{"task_slug":"hallucination","task_name":"Hallucination"},{"task_slug":"optical-flow-estimation","task_name":"Optical Flow Estimation"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1712.04109","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}