{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/iqiyi-vid-a-large-dataset-for-multi-modal","title":"iQIYI-VID: A Large Dataset for Multi-modal Person Identification","arxiv_id":"1811.07548","date":"2018-11-19","proceeding":null,"authors":["Yuanliu Liu","Bo Peng","Peipei Shi","He Yan","Yong Zhou","Bing Han","Yi Zheng","Chao Lin","Jianbin Jiang","Yin Fan","Tingwei Gao","Ganwen Wang","Jian Liu","Xiangju Lu","Danming Xie"],"abstract":"Person identification in the wild is very challenging due to great variation\nin poses, face quality, clothes, makeup and so on. Traditional research, such\nas face recognition, person re-identification, and speaker recognition, often\nfocuses on a single modal of information, which is inadequate to handle all the\nsituations in practice. Multi-modal person identification is a more promising\nway that we can jointly utilize face, head, body, audio features, and so on. In\nthis paper, we introduce iQIYI-VID, the largest video dataset for multi-modal\nperson identification. It is composed of 600K video clips of 5,000 celebrities.\nThese video clips are extracted from 400K hours of online videos of various\ntypes, ranging from movies, variety shows, TV series, to news broadcasting. All\nvideo clips pass through a careful human annotation process, and the error rate\nof labels is lower than 0.2\\%. We evaluated the state-of-art models of face\nrecognition, person re-identification, and speaker recognition on the iQIYI-VID\ndataset. Experimental results show that these models are still far from being\nperfect for the task of person identification in the wild. We proposed a\nMulti-modal Attention module to fuse multi-modal features that can improve\nperson identification considerably. We have released the dataset online to\npromote multi-modal person identification research.","url_abs":"http://arxiv.org/abs/1811.07548v2","url_pdf":"http://arxiv.org/pdf/1811.07548v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"face-recognition","task_name":"Face Recognition"},{"task_slug":"multi-modal-person-identification","task_name":"Multi-Modal Person Identification"},{"task_slug":"person-identification","task_name":"Person Identification"},{"task_slug":"person-re-identification","task_name":"Person Re-Identification"},{"task_slug":"speaker-recognition","task_name":"Speaker Recognition"}],"methods":[],"datasets_introduced":[{"slug":"iqiyi-vid","name":"iQIYI-VID","full_name":"iQIYI-VID"}],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1811.07548","atlas_url":"https://app.syntology.ai/?focus=1811.07548","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}