{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/disentangled-person-image-generation","title":"Disentangled Person Image Generation","arxiv_id":"1712.02621","date":"2017-12-07","proceeding":"CVPR 2018 6","authors":["Liqian Ma","Qianru Sun","Stamatios Georgoulis","Luc van Gool","Bernt Schiele","Mario Fritz"],"abstract":"Generating novel, yet realistic, images of persons is a challenging task due\nto the complex interplay between the different image factors, such as the\nforeground, background and pose information. In this work, we aim at generating\nsuch images based on a novel, two-stage reconstruction pipeline that learns a\ndisentangled representation of the aforementioned image factors and generates\nnovel person images at the same time. First, a multi-branched reconstruction\nnetwork is proposed to disentangle and encode the three factors into embedding\nfeatures, which are then combined to re-compose the input image itself. Second,\nthree corresponding mapping functions are learned in an adversarial manner in\norder to map Gaussian noise to the learned embedding feature space, for each\nfactor respectively. Using the proposed framework, we can manipulate the\nforeground, background and pose of the input image, and also sample new\nembedding features to generate such targeted manipulations, that provide more\ncontrol over the generation process. Experiments on Market-1501 and Deepfashion\ndatasets show that our model does not only generate realistic person images\nwith new foregrounds, backgrounds and poses, but also manipulates the generated\nfactors and interpolates the in-between states. Another set of experiments on\nMarket-1501 shows that our model can also be beneficial for the person\nre-identification task.","url_abs":"http://arxiv.org/abs/1712.02621v4","url_pdf":"http://arxiv.org/pdf/1712.02621v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"disentangled-person-image-generation","repo_url":"https://github.com/charliememory/Disentangled-Person-Image-Generation","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"tf","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"gesture-to-gesture-translation","task_name":"Gesture-to-Gesture Translation"},{"task_slug":"image-generation","task_name":"Image Generation"},{"task_slug":"person-re-identification","task_name":"Person Re-Identification"},{"task_slug":"pose-transfer","task_name":"Pose Transfer"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/gesture-to-gesture-translation-on-ntu-hand","task":"Gesture-to-Gesture Translation","dataset":"NTU Hand Digit","model":"DPIG","rank_in_archive_order":3,"of":6,"metrics":{"AMT":"7.1","IS":"2.4547","PSNR":"30.6487"},"uses_additional_data":false},{"leaderboard":"/sota/gesture-to-gesture-translation-on-senz3d","task":"Gesture-to-Gesture Translation","dataset":"Senz3D","model":"DPIG","rank_in_archive_order":2,"of":6,"metrics":{"AMT":"6.9","IS":"3.3874","PSNR":"26.9451"},"uses_additional_data":false},{"leaderboard":"/sota/pose-transfer-on-deep-fashion","task":"Pose Transfer","dataset":"Deep-Fashion","model":"Disentangled PG","rank_in_archive_order":10,"of":12,"metrics":{"IS":"3.228","SSIM":"0.614"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/1712.02621","atlas_url":"https://app.syntology.ai/?focus=1712.02621","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}