{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/gen3c-3d-informed-world-consistent-video","title":"GEN3C: 3D-Informed World-Consistent Video Generation with Precise Camera Control","arxiv_id":"2503.03751","date":"2025-03-05","proceeding":"CVPR 2025 1","authors":["Xuanchi Ren","Tianchang Shen","Jiahui Huang","Huan Ling","Yifan Lu","Merlin Nimier-David","Thomas Müller","Alexander Keller","Sanja Fidler","Jun Gao"],"abstract":"We present GEN3C, a generative video model with precise Camera Control and temporal 3D Consistency. Prior video models already generate realistic videos, but they tend to leverage little 3D information, leading to inconsistencies, such as objects popping in and out of existence. Camera control, if implemented at all, is imprecise, because camera parameters are mere inputs to the neural network which must then infer how the video depends on the camera. In contrast, GEN3C is guided by a 3D cache: point clouds obtained by predicting the pixel-wise depth of seed images or previously generated frames. When generating the next frames, GEN3C is conditioned on the 2D renderings of the 3D cache with the new camera trajectory provided by the user. Crucially, this means that GEN3C neither has to remember what it previously generated nor does it have to infer the image structure from the camera pose. The model, instead, can focus all its generative power on previously unobserved regions, as well as advancing the scene state to the next frame. Our results demonstrate more precise camera control than prior work, as well as state-of-the-art results in sparse-view novel view synthesis, even in challenging settings such as driving scenes and monocular dynamic video. Results are best viewed in videos. Check out our webpage! https://research.nvidia.com/labs/toronto-ai/GEN3C/","url_abs":"https://arxiv.org/abs/2503.03751v1","url_pdf":"https://arxiv.org/pdf/2503.03751v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"gen3c-3d-informed-world-consistent-video","repo_url":"https://github.com/nv-tlabs/GEN3C","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":"novel-view-synthesis","task_name":"Novel View Synthesis"},{"task_slug":"video-generation","task_name":"Video Generation"}],"methods":[{"method_slug":"focus","method_name":"Focus"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2503.03751","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2503.03751"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/nv-tlabs/GEN3C","reach":{"status":"ok","spdx":"Apache-2.0"}}],"summary":{"unverified":2},"by_repo_kind":{"official":{"samples":2,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"25b0f02063e9d156","entry":"resize_intrinsics","repo":"nv-tlabs/GEN3C","repo_kind":"official","path":"cosmos_predict1/diffusion/inference/gen3c_persistent.py","file_url":"https://github.com/nv-tlabs/GEN3C/blob/HEAD/cosmos_predict1/diffusion/inference/gen3c_persistent.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"25b0f02063e9d156"}},{"code_sha256_prefix":"abc31b2f673701c7","entry":"scaled_dot_product_attention","repo":"nv-tlabs/GEN3C","repo_kind":"official","path":"cosmos_predict1/autoregressive/modules/attention.py","file_url":"https://github.com/nv-tlabs/GEN3C/blob/HEAD/cosmos_predict1/autoregressive/modules/attention.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"abc31b2f673701c7"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}