{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/tracking-persons-of-interest-via-unsupervised","title":"Tracking Persons-of-Interest via Unsupervised Representation Adaptation","arxiv_id":"1710.02139","date":"2017-10-05","proceeding":null,"authors":["Shun Zhang","Jia-Bin Huang","Jongwoo Lim","Yihong Gong","Jinjun Wang","Narendra Ahuja","Ming-Hsuan Yang"],"abstract":"Multi-face tracking in unconstrained videos is a challenging problem as faces\nof one person often appear drastically different in multiple shots due to\nsignificant variations in scale, pose, expression, illumination, and make-up.\nExisting multi-target tracking methods often use low-level features which are\nnot sufficiently discriminative for identifying faces with such large\nappearance variations. In this paper, we tackle this problem by learning\ndiscriminative, video-specific face representations using convolutional neural\nnetworks (CNNs). Unlike existing CNN-based approaches which are only trained on\nlarge-scale face image datasets offline, we use the contextual constraints to\ngenerate a large number of training samples for a given video, and further\nadapt the pre-trained face CNN to specific videos using discovered training\nsamples. Using these training samples, we optimize the embedding space so that\nthe Euclidean distances correspond to a measure of semantic face similarity via\nminimizing a triplet loss function. With the learned discriminative features,\nwe apply the hierarchical clustering algorithm to link tracklets across\nmultiple shots to generate trajectories. We extensively evaluate the proposed\nalgorithm on two sets of TV sitcoms and YouTube music videos, analyze the\ncontribution of each component, and demonstrate significant performance\nimprovement over existing techniques.","url_abs":"http://arxiv.org/abs/1710.02139v1","url_pdf":"http://arxiv.org/pdf/1710.02139v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"tracking-persons-of-interest-via-unsupervised","repo_url":"https://github.com/shunzhang876/AdaptiveFeatureLearning","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"none","reach":null},{"paper_slug":"tracking-persons-of-interest-via-unsupervised","repo_url":"https://github.com/noelcodella/triplet_loss_pytorch","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"clustering","task_name":"Clustering"},{"task_slug":null,"task_name":"Triplet"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}