Papers › Text-guided 3D Human Generation from 2D Collections

Text-guided 3D Human Generation from 2D Collections

23 May 2023arXiv:2305.14312archive 2025-07-28

Tsu-Jui Fu, Wenhan Xiong, Yixin Nie, Jingyu Liu, Barlas Oğuz, William Yang Wang

3D human modeling has been widely used for engaging interaction in gaming, film, and animation. The customization of these characters is crucial for creativity and scalability, which highlights the importance of controllability. In this work, we introduce Text-guided 3D Human Generation (\texttt{T3H}), where a model is to generate a 3D human, guided by the fashion description. There are two goals: 1) the 3D human should render articulately, and 2) its outfit is controlled by the given text. To address this \texttt{T3H} task, we propose Compositional Cross-modal Human (CCH). CCH adopts cross-modal attention to fuse compositional human rendering with the extracted fashion semantics. Each human body part perceives relevant textual guidance as its visual patterns. We incorporate the human prior and semantic discrimination to enhance 3D geometry transformation and fine-grained consistency, enabling it to learn from 2D collections for data efficiency. We conduct evaluations on DeepFashion and SHHQ with diverse fashion attributes covering the shape, fabric, and color of upper and lower clothing. Extensive experiments demonstrate that CCH achieves superior results for \texttt{T3H} with high efficiency.

PaperPDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

3D geometryText-to-3D-Human Generationtext-to-3d-human

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Text-to-3D-Human Generation DeepFashion CCH CLIP Score 25.031 #1 of 1 Archive leaderboard report
Text-to-3D-Human Generation DeepFashion CCH Depth Error 1.21 #1 of 1 Archive leaderboard report
Text-to-3D-Human Generation DeepFashion CCH Fashion Accuracy 72.038 #1 of 1 Archive leaderboard report
Text-to-3D-Human Generation DeepFashion CCH Frechet Inception Distance 22.175 #1 of 1 Archive leaderboard report
Text-to-3D-Human Generation DeepFashion CCH Percentage of Correct Keypoints 88.313 #1 of 1 Archive leaderboard report
Text-to-3D-Human Generation SHHQ CCH CLIP Score 27.873 #1 of 1 Archive leaderboard report
Text-to-3D-Human Generation SHHQ CCH Depth Error 1.67 #1 of 1 Archive leaderboard report
Text-to-3D-Human Generation SHHQ CCH Fashion Accuracy 76.194 #1 of 1 Archive leaderboard report
Text-to-3D-Human Generation SHHQ CCH Frechet Inception Distance 33.348 #1 of 1 Archive leaderboard report
Text-to-3D-Human Generation SHHQ CCH Percentage of Correct Keypoints 87.879 #1 of 1 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections