{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/clustering-millions-of-faces-by-identity","title":"Clustering Millions of Faces by Identity","arxiv_id":"1604.00989","date":"2016-04-04","proceeding":null,"authors":["Charles Otto","Dayong Wang","Anil K. Jain"],"abstract":"In this work, we attempt to address the following problem: Given a large\nnumber of unlabeled face images, cluster them into the individual identities\npresent in this data. We consider this a relevant problem in different\napplication scenarios ranging from social media to law enforcement. In\nlarge-scale scenarios the number of faces in the collection can be of the order\nof hundreds of million, while the number of clusters can range from a few\nthousand to millions--leading to difficulties in terms of both run-time\ncomplexity and evaluating clustering and per-cluster quality. An efficient and\neffective Rank-Order clustering algorithm is developed to achieve the desired\nscalability, and better clustering accuracy than other well-known algorithms\nsuch as k-means and spectral clustering. We cluster up to 123 million face\nimages into over 10 million clusters, and analyze the results in terms of both\nexternal cluster quality measures (known face labels) and internal cluster\nquality measures (unknown face labels) and run-time. Our algorithm achieves an\nF-measure of 0.87 on a benchmark unconstrained face dataset (LFW, consisting of\n13K faces), and 0.27 on the largest dataset considered (13K images in LFW, plus\n123M distractor images). Additionally, we present preliminary work on video\nframe clustering (achieving 0.71 F-measure when clustering all frames in the\nbenchmark YouTube Faces dataset). A per-cluster quality measure is developed\nwhich can be used to rank individual clusters and to automatically identify a\nsubset of good quality clusters for manual exploration.","url_abs":"http://arxiv.org/abs/1604.00989v1","url_pdf":"http://arxiv.org/pdf/1604.00989v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"clustering-millions-of-faces-by-identity","repo_url":"https://github.com/varun-suresh/Clustering","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"clustering","task_name":"Clustering"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1604.00989","atlas_url":"https://app.syntology.ai/?focus=1604.00989","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}