{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/distractor-aware-siamese-networks-for-visual","title":"Distractor-aware Siamese Networks for Visual Object Tracking","arxiv_id":"1808.06048","date":"2018-08-18","proceeding":"ECCV 2018 9","authors":["Zheng Zhu","Qiang Wang","Bo Li","Wei Wu","Junjie Yan","Weiming Hu"],"abstract":"Recently, Siamese networks have drawn great attention in visual tracking\ncommunity because of their balanced accuracy and speed. However, features used\nin most Siamese tracking approaches can only discriminate foreground from the\nnon-semantic backgrounds. The semantic backgrounds are always considered as\ndistractors, which hinders the robustness of Siamese trackers. In this paper,\nwe focus on learning distractor-aware Siamese networks for accurate and\nlong-term tracking. To this end, features used in traditional Siamese trackers\nare analyzed at first. We observe that the imbalanced distribution of training\ndata makes the learned features less discriminative. During the off-line\ntraining phase, an effective sampling strategy is introduced to control this\ndistribution and make the model focus on the semantic distractors. During\ninference, a novel distractor-aware module is designed to perform incremental\nlearning, which can effectively transfer the general embedding to the current\nvideo domain. In addition, we extend the proposed approach for long-term\ntracking by introducing a simple yet effective local-to-global search region\nstrategy. Extensive experiments on benchmarks show that our approach\nsignificantly outperforms the state-of-the-arts, yielding 9.6% relative gain in\nVOT2016 dataset and 35.9% relative gain in UAV20L dataset. The proposed tracker\ncan perform at 160 FPS on short-term benchmarks and 110 FPS on long-term\nbenchmarks.","url_abs":"http://arxiv.org/abs/1808.06048v1","url_pdf":"http://arxiv.org/pdf/1808.06048v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"distractor-aware-siamese-networks-for-visual","repo_url":"https://github.com/foolwood/DaSiamRPN","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"incremental-learning","task_name":"Incremental Learning"},{"task_slug":"object","task_name":"Object"},{"task_slug":"object-tracking","task_name":"Object Tracking"},{"task_slug":"video-object-tracking","task_name":"Video Object Tracking"},{"task_slug":"visual-object-tracking","task_name":"Visual Object Tracking"},{"task_slug":"visual-tracking","task_name":"Visual Tracking"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/video-object-tracking-on-nv-vot211","task":"Video Object Tracking","dataset":"NT-VOT211","model":"DaSiamRPN","rank_in_archive_order":32,"of":43,"metrics":{"AUC":"31.12","Precision":"39.09"},"uses_additional_data":false},{"leaderboard":"/sota/visual-object-tracking-on-vot201718","task":"Visual Object Tracking","dataset":"VOT2017/18","model":"DaSiamRPN","rank_in_archive_order":11,"of":15,"metrics":{"Expected Average Overlap (EAO)":"0.326"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/1808.06048","atlas_url":"https://app.syntology.ai/?focus=1808.06048","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}