{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/from-an-image-to-a-scene-learning-to-imagine","title":"From an Image to a Scene: Learning to Imagine the World from a Million 360 Videos","arxiv_id":"2412.07770","date":"2024-12-10","proceeding":null,"authors":["Matthew Wallingford","Anand Bhattad","Aditya Kusupati","Vivek Ramanujan","Matt Deitke","Sham Kakade","Aniruddha Kembhavi","Roozbeh Mottaghi","Wei-Chiu Ma","Ali Farhadi"],"abstract":"Three-dimensional (3D) understanding of objects and scenes play a key role in humans' ability to interact with the world and has been an active area of research in computer vision, graphics, and robotics. Large scale synthetic and object-centric 3D datasets have shown to be effective in training models that have 3D understanding of objects. However, applying a similar approach to real-world objects and scenes is difficult due to a lack of large-scale data. Videos are a potential source for real-world 3D data, but finding diverse yet corresponding views of the same content has shown to be difficult at scale. Furthermore, standard videos come with fixed viewpoints, determined at the time of capture. This restricts the ability to access scenes from a variety of more diverse and potentially useful perspectives. We argue that large scale 360 videos can address these limitations to provide: scalable corresponding frames from diverse views. In this paper, we introduce 360-1M, a 360 video dataset, and a process for efficiently finding corresponding frames from diverse viewpoints at scale. We train our diffusion-based model, Odin, on 360-1M. Empowered by the largest real-world, multi-view dataset to date, Odin is able to freely generate novel views of real-world scenes. Unlike previous methods, Odin can move the camera through the environment, enabling the model to infer the geometry and layout of the scene. Additionally, we show improved performance on standard novel view synthesis and 3D reconstruction benchmarks.","url_abs":"https://arxiv.org/abs/2412.07770v1","url_pdf":"https://arxiv.org/pdf/2412.07770v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"from-an-image-to-a-scene-learning-to-imagine","repo_url":"https://github.com/mattwallingford/360-1m","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"3d-reconstruction","task_name":"3D Reconstruction"},{"task_slug":"novel-view-synthesis","task_name":"Novel View Synthesis"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2412.07770","atlas_url":"https://app.syntology.ai/?focus=2412.07770","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2412.07770"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/mattwallingford/360-1m","reach":{"status":"ok"}}],"summary":{"unverified":3},"by_repo_kind":{"official":{"samples":3,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":3,"samples":[{"code_sha256_prefix":"d07895a2f6ed02f8","entry":"filter_pairs_seq","repo":"mattwallingford/360-1m","repo_kind":"official","path":"VideoProcessing/extract_poses.py","file_url":"https://github.com/mattwallingford/360-1m/blob/HEAD/VideoProcessing/extract_poses.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"d07895a2f6ed02f8"}},{"code_sha256_prefix":"c4bbf32f86a5dec1","entry":"make_pairs","repo":"mattwallingford/360-1m","repo_kind":"official","path":"VideoProcessing/extract_poses.py","file_url":"https://github.com/mattwallingford/360-1m/blob/HEAD/VideoProcessing/extract_poses.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"c4bbf32f86a5dec1"}},{"code_sha256_prefix":"88e89104d153a472","entry":"sel","repo":"mattwallingford/360-1m","repo_kind":"official","path":"VideoProcessing/extract_poses.py","file_url":"https://github.com/mattwallingford/360-1m/blob/HEAD/VideoProcessing/extract_poses.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"88e89104d153a472"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}