{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/dynamic-computational-time-for-visual","title":"Dynamic Computational Time for Visual Attention","arxiv_id":"1703.10332","date":"2017-03-30","proceeding":null,"authors":["Zhichao Li","Yi Yang","Xiao Liu","Feng Zhou","Shilei Wen","Wei Xu"],"abstract":"We propose a dynamic computational time model to accelerate the average\nprocessing time for recurrent visual attention (RAM). Rather than attention\nwith a fixed number of steps for each input image, the model learns to decide\nwhen to stop on the fly. To achieve this, we add an additional continue/stop\naction per time step to RAM and use reinforcement learning to learn both the\noptimal attention policy and stopping policy. The modification is simple but\ncould dramatically save the average computational time while keeping the same\nrecognition performance as RAM. Experimental results on CUB-200-2011 and\nStanford Cars dataset demonstrate the dynamic computational model can work\neffectively for fine-grained image recognition.The source code of this paper\ncan be obtained from https://github.com/baidu-research/DT-RAM","url_abs":"http://arxiv.org/abs/1703.10332v3","url_pdf":"http://arxiv.org/pdf/1703.10332v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"dynamic-computational-time-for-visual","repo_url":"https://github.com/baidu-research/DT-RAM","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"torch","reach":{"status":"ok"}},{"paper_slug":"dynamic-computational-time-for-visual","repo_url":"https://github.com/kiwi1944/CRISense","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1703.10332","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}