{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/towards-high-performance-video-object","title":"Towards High Performance Video Object Detection for Mobiles","arxiv_id":"1804.05830","date":"2018-04-16","proceeding":null,"authors":["Xizhou Zhu","Jifeng Dai","Xingchi Zhu","Yichen Wei","Lu Yuan"],"abstract":"Despite the recent success of video object detection on Desktop GPUs, its\narchitecture is still far too heavy for mobiles. It is also unclear whether the\nkey principles of sparse feature propagation and multi-frame feature\naggregation apply at very limited computational resources. In this paper, we\npresent a light weight network architecture for video object detection on\nmobiles. Light weight image object detector is applied on sparse key frames. A\nvery small network, Light Flow, is designed for establishing correspondence\nacross frames. A flow-guided GRU module is designed to effectively aggregate\nfeatures on key frames. For non-key frames, sparse feature propagation is\nperformed. The whole network can be trained end-to-end. The proposed system\nachieves 60.2% mAP score at speed of 25.6 fps on mobiles (e.g., HuaWei Mate 8).","url_abs":"http://arxiv.org/abs/1804.05830v1","url_pdf":"http://arxiv.org/pdf/1804.05830v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"towards-high-performance-video-object","repo_url":"https://github.com/McDo/LightFlowPytorch","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"towards-high-performance-video-object","repo_url":"https://github.com/stanlee321/LightFlow-Keras","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null},{"paper_slug":"towards-high-performance-video-object","repo_url":"https://github.com/stanlee321/LightFlow-TensorFlow","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"object","task_name":"Object"},{"task_slug":"object-detection","task_name":"Object Detection"},{"task_slug":"video-object-detection","task_name":"Video Object Detection"},{"task_slug":"high","task_name":"Vocal Bursts Intensity Prediction"},{"task_slug":"object-detection-1","task_name":"object-detection"}],"methods":[{"method_slug":"gru","method_name":"GRU"},{"method_slug":"speed","method_name":"SPEED"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1804.05830","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}