{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/dual-deep-network-for-visual-tracking","title":"Dual Deep Network for Visual Tracking","arxiv_id":"1612.06053","date":"2016-12-19","proceeding":null,"authors":["Zhizhen Chi","Hongyang Li","Huchuan Lu","Ming-Hsuan Yang"],"abstract":"Visual tracking addresses the problem of identifying and localizing an\nunknown target in a video given the target specified by a bounding box in the\nfirst frame. In this paper, we propose a dual network to better utilize\nfeatures among layers for visual tracking. It is observed that features in\nhigher layers encode semantic context while its counterparts in lower layers\nare sensitive to discriminative appearance. Thus we exploit the hierarchical\nfeatures in different layers of a deep model and design a dual structure to\nobtain better feature representation from various streams, which is rarely\ninvestigated in previous work. To highlight geometric contours of the target,\nwe integrate the hierarchical feature maps with an edge detector as the coarse\nprior maps to further embed local details around the target. To leverage the\nrobustness of our dual network, we train it with random patches measuring the\nsimilarities between the network activation and target appearance, which serves\nas a regularization to enforce the dual network to focus on target object. The\nproposed dual network is updated online in a unique manner based on the\nobservation that the target being tracked in consecutive frames should share\nmore similar feature representations than those in the surrounding background.\nIt is also found that for a target object, the prior maps can help further\nenhance performance by passing message into the output maps of the dual\nnetwork. Therefore, an independent component analysis with reference algorithm\n(ICA-R) is employed to extract target context using prior maps as guidance.\nOnline tracking is conducted by maximizing the posterior estimate on the final\nmaps with stochastic and periodic update. Quantitative and qualitative\nevaluations on two large-scale benchmark data sets show that the proposed\nalgorithm performs favourably against the state-of-the-arts.","url_abs":"http://arxiv.org/abs/1612.06053v1","url_pdf":"http://arxiv.org/pdf/1612.06053v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"dual-deep-network-for-visual-tracking","repo_url":"https://github.com/chizhizhen/DNT","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"visual-tracking","task_name":"Visual Tracking"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}