{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-multi-domain-convolutional-neural","title":"Learning Multi-Domain Convolutional Neural Networks for Visual Tracking","arxiv_id":"1510.07945","date":"2015-10-27","proceeding":"CVPR 2016 6","authors":["Hyeonseob Nam","Bohyung Han"],"abstract":"We propose a novel visual tracking algorithm based on the representations\nfrom a discriminatively trained Convolutional Neural Network (CNN). Our\nalgorithm pretrains a CNN using a large set of videos with tracking\nground-truths to obtain a generic target representation. Our network is\ncomposed of shared layers and multiple branches of domain-specific layers,\nwhere domains correspond to individual training sequences and each branch is\nresponsible for binary classification to identify the target in each domain. We\ntrain the network with respect to each domain iteratively to obtain generic\ntarget representations in the shared layers. When tracking a target in a new\nsequence, we construct a new network by combining the shared layers in the\npretrained CNN with a new binary classification layer, which is updated online.\nOnline tracking is performed by evaluating the candidate windows randomly\nsampled around the previous target state. The proposed algorithm illustrates\noutstanding performance compared with state-of-the-art methods in existing\ntracking benchmarks.","url_abs":"http://arxiv.org/abs/1510.07945v2","url_pdf":"http://arxiv.org/pdf/1510.07945v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"learning-multi-domain-convolutional-neural","repo_url":"https://github.com/FelixOliver/PROYECTO-FINAL-VOT","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}},{"paper_slug":"learning-multi-domain-convolutional-neural","repo_url":"https://github.com/HyeonseobNam/py-MDNet","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"NOASSERTION"}}],"tasks":[{"task_slug":"binary-classification","task_name":"Binary Classification"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"visual-tracking","task_name":"Visual Tracking"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/visual-tracking-on-second-dialogue-state","task":"Visual Tracking","dataset":"Second dialogue state tracking challenge","model":"MDNet","rank_in_archive_order":1,"of":1,"metrics":{"Score":"0.64"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/1510.07945","atlas_url":"https://app.syntology.ai/?focus=1510.07945","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}