{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/pifu-pixel-aligned-implicit-function-for-high","title":"PIFu: Pixel-Aligned Implicit Function for High-Resolution Clothed Human Digitization","arxiv_id":"1905.05172","date":"2019-05-13","proceeding":"ICCV 2019 10","authors":["Shunsuke Saito","Zeng Huang","Ryota Natsume","Shigeo Morishima","Angjoo Kanazawa","Hao Li"],"abstract":"We introduce Pixel-aligned Implicit Function (PIFu), a highly effective implicit representation that locally aligns pixels of 2D images with the global context of their corresponding 3D object. Using PIFu, we propose an end-to-end deep learning method for digitizing highly detailed clothed humans that can infer both 3D surface and texture from a single image, and optionally, multiple input images. Highly intricate shapes, such as hairstyles, clothing, as well as their variations and deformations can be digitized in a unified way. Compared to existing representations used for 3D deep learning, PIFu can produce high-resolution surfaces including largely unseen regions such as the back of a person. In particular, it is memory efficient unlike the voxel representation, can handle arbitrary topology, and the resulting surface is spatially aligned with the input image. Furthermore, while previous techniques are designed to process either a single image or multiple views, PIFu extends naturally to arbitrary number of views. We demonstrate high-resolution and robust reconstructions on real world images from the DeepFashion dataset, which contains a variety of challenging clothing types. Our method achieves state-of-the-art performance on a public benchmark and outperforms the prior work for clothed human digitization from a single image.","url_abs":"https://arxiv.org/abs/1905.05172v3","url_pdf":"https://arxiv.org/pdf/1905.05172v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"pifu-pixel-aligned-implicit-function-for-high","repo_url":"https://github.com/shunsukesaito/PIFu","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"NOASSERTION"}}],"tasks":[{"task_slug":"3d-human-pose-estimation","task_name":"3D Human Pose Estimation"},{"task_slug":"3d-human-reconstruction","task_name":"3D Human Reconstruction"},{"task_slug":"3d-object-reconstruction","task_name":"3D Object Reconstruction"},{"task_slug":"3d-object-reconstruction-from-a-single-image","task_name":"3D Object Reconstruction From A Single Image"},{"task_slug":"3d-shape-reconstruction","task_name":"3D Shape Reconstruction"},{"task_slug":"3d-shape-reconstruction-from-a-single-2d","task_name":"3D Shape Reconstruction From A Single 2D Image"},{"task_slug":"lifelike-3d-human-generation","task_name":"Lifelike 3D Human Generation"},{"task_slug":"high","task_name":"Vocal Bursts Intensity Prediction"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/3d-human-reconstruction-on-4d-dress","task":"3D Human Reconstruction","dataset":"4D-DRESS","model":"PIFu_Inner","rank_in_archive_order":15,"of":22,"metrics":{"Chamfer (cm)":"2.696","IoU":"0.690","Normal Consistency":"0.792"},"uses_additional_data":true},{"leaderboard":"/sota/3d-human-reconstruction-on-4d-dress","task":"3D Human Reconstruction","dataset":"4D-DRESS","model":"PIFu_Outer","rank_in_archive_order":16,"of":22,"metrics":{"Chamfer (cm)":"2.783","IoU":"0.697","Normal Consistency":"0.759"},"uses_additional_data":true},{"leaderboard":"/sota/3d-human-reconstruction-on-cape","task":"3D Human Reconstruction","dataset":"CAPE","model":"PIFu (THuman2.0)","rank_in_archive_order":4,"of":4,"metrics":{"Chamfer (cm)":"3.573","NC":"0.186","P2S (cm)":"1.483"},"uses_additional_data":true},{"leaderboard":"/sota/3d-human-reconstruction-on-customhumans","task":"3D Human Reconstruction","dataset":"CustomHumans","model":"PIFu","rank_in_archive_order":5,"of":8,"metrics":{"Chamfer Distance P-to-S":"2.209","Chamfer Distance S-to-P":"2.582","Normal Consistency":"0.805","f-Score":"34.881"},"uses_additional_data":true},{"leaderboard":"/sota/3d-object-reconstruction-on-renderpeople","task":"3D Object Reconstruction","dataset":"RenderPeople","model":"PIFu (3 views)","rank_in_archive_order":1,"of":1,"metrics":{"Chamfer (cm)":"0.567","Point-to-surface distance (cm)":"0.554","Surface normal consistency":"0.094"},"uses_additional_data":false},{"leaderboard":"/sota/3d-object-reconstruction-from-a-single-image-1","task":"3D Object Reconstruction From A Single Image","dataset":"BUFF","model":"PIFu","rank_in_archive_order":4,"of":5,"metrics":{"Chamfer (cm)":"1.14","Point-to-surface distance (cm)":"1.15","Surface normal consistency":"0.0928"},"uses_additional_data":false},{"leaderboard":"/sota/3d-object-reconstruction-from-a-single-image","task":"3D Object Reconstruction From A Single Image","dataset":"RenderPeople","model":"PIFu","rank_in_archive_order":2,"of":4,"metrics":{"Chamfer (cm)":"1.5","Point-to-surface distance (cm)":"1.52","Surface normal consistency":"0.084"},"uses_additional_data":true},{"leaderboard":"/sota/lifelike-3d-human-generation-on-thuman2-0","task":"Lifelike 3D Human Generation","dataset":"THuman2.0 Dataset","model":"PIFu","rank_in_archive_order":6,"of":6,"metrics":{"CLIP Similarity":"0.8501","LPIPS":"0.1615","PSNR":"15.0248","SSIM":"0.8884"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1905.05172","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}