{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/neural-aggregation-network-for-video-face","title":"Neural Aggregation Network for Video Face Recognition","arxiv_id":"1603.05474","date":"2016-03-17","proceeding":"CVPR 2017 7","authors":["Jiaolong Yang","Peiran Ren","Dong-Qing Zhang","Dong Chen","Fang Wen","Hongdong Li","Gang Hua"],"abstract":"This paper presents a Neural Aggregation Network (NAN) for video face\nrecognition. The network takes a face video or face image set of a person with\na variable number of face images as its input, and produces a compact,\nfixed-dimension feature representation for recognition. The whole network is\ncomposed of two modules. The feature embedding module is a deep Convolutional\nNeural Network (CNN) which maps each face image to a feature vector. The\naggregation module consists of two attention blocks which adaptively aggregate\nthe feature vectors to form a single feature inside the convex hull spanned by\nthem. Due to the attention mechanism, the aggregation is invariant to the image\norder. Our NAN is trained with a standard classification or verification loss\nwithout any extra supervision signal, and we found that it automatically learns\nto advocate high-quality face images while repelling low-quality ones such as\nblurred, occluded and improperly exposed faces. The experiments on IJB-A,\nYouTube Face, Celebrity-1000 video face recognition benchmarks show that it\nconsistently outperforms naive aggregation methods and achieves the\nstate-of-the-art accuracy.","url_abs":"http://arxiv.org/abs/1603.05474v4","url_pdf":"http://arxiv.org/pdf/1603.05474v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"face-identification","task_name":"Face Identification"},{"task_slug":"face-recognition","task_name":"Face Recognition"},{"task_slug":"face-verification","task_name":"Face Verification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/face-identification-on-dronesurf","task":"Face Identification","dataset":"DroneSURF","model":"NAN (Adaface)","rank_in_archive_order":3,"of":6,"metrics":{"Rank1":"80.21"},"uses_additional_data":false},{"leaderboard":"/sota/face-verification-on-bts3-1","task":"Face Verification","dataset":"BTS3.1","model":"NAN (Adaface)","rank_in_archive_order":3,"of":7,"metrics":{"TAR @ FAR=0.01":"0.5444"},"uses_additional_data":false},{"leaderboard":"/sota/face-verification-on-bts3-1","task":"Face Verification","dataset":"BTS3.1","model":"MCN (Arcface)","rank_in_archive_order":6,"of":7,"metrics":{"TAR @ FAR=0.01":"0.3941"},"uses_additional_data":false},{"leaderboard":"/sota/face-verification-on-bts3-1","task":"Face Verification","dataset":"BTS3.1","model":"NAN (Arcface)","rank_in_archive_order":7,"of":7,"metrics":{"TAR @ FAR=0.01":"0.3901"},"uses_additional_data":false},{"leaderboard":"/sota/face-verification-on-ijb-a","task":"Face Verification","dataset":"IJB-A","model":"NAN","rank_in_archive_order":8,"of":17,"metrics":{"TAR @ FAR=0.01":"94.10%"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1603.05474","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}