{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/kernalised-multi-resolution-convnet-for","title":"Kernalised Multi-resolution Convnet for Visual Tracking","arxiv_id":"1708.00577","date":"2017-08-02","proceeding":null,"authors":["Di Wu","Wenbin Zou","Xia Li","Yong Zhao"],"abstract":"Visual tracking is intrinsically a temporal problem. Discriminative\nCorrelation Filters (DCF) have demonstrated excellent performance for\nhigh-speed generic visual object tracking. Built upon their seminal work, there\nhas been a plethora of recent improvements relying on convolutional neural\nnetwork (CNN) pretrained on ImageNet as a feature extractor for visual\ntracking. However, most of their works relying on ad hoc analysis to design the\nweights for different layers either using boosting or hedging techniques as an\nensemble tracker. In this paper, we go beyond the conventional DCF framework\nand propose a Kernalised Multi-resolution Convnet (KMC) formulation that\nutilises hierarchical response maps to directly output the target movement.\nWhen directly deployed the learnt network to predict the unseen challenging UAV\ntracking dataset without any weight adjustment, the proposed model consistently\nachieves excellent tracking performance. Moreover, the transfered\nmulti-reslution CNN renders it possible to be integrated into the RNN temporal\nlearning framework, therefore opening the door on the end-to-end temporal deep\nlearning (TDL) for visual tracking.","url_abs":"http://arxiv.org/abs/1708.00577v1","url_pdf":"http://arxiv.org/pdf/1708.00577v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"kernalised-multi-resolution-convnet-for","repo_url":"https://github.com/stevenwudi/KMC_cvprw_2017","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"tf","reach":null}],"tasks":[{"task_slug":"object-tracking","task_name":"Object Tracking"},{"task_slug":"visual-object-tracking","task_name":"Visual Object Tracking"},{"task_slug":"visual-tracking","task_name":"Visual Tracking"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}