{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-to-predict-without-looking-ahead","title":"Learning to Predict Without Looking Ahead: World Models Without Forward Prediction","arxiv_id":"1910.13038","date":"2019-10-29","proceeding":"NeurIPS 2019 12","authors":["C. Daniel Freeman","Luke Metz","David Ha"],"abstract":"Much of model-based reinforcement learning involves learning a model of an agent's world, and training an agent to leverage this model to perform a task more efficiently. While these models are demonstrably useful for agents, every naturally occurring model of the world of which we are aware---e.g., a brain---arose as the byproduct of competing evolutionary pressures for survival, not minimization of a supervised forward-predictive loss via gradient descent. That useful models can arise out of the messy and slow optimization process of evolution suggests that forward-predictive modeling can arise as a side-effect of optimization under the right circumstances. Crucially, this optimization process need not explicitly be a forward-predictive loss. In this work, we introduce a modification to traditional reinforcement learning which we call observational dropout, whereby we limit the agents ability to observe the real environment at each timestep. In doing so, we can coerce an agent into learning a world model to fill in the observation gaps during reinforcement learning. We show that the emerged world model, while not explicitly trained to predict the future, can help the agent learn key skills required to perform well in its environment. Videos of our results available at https://learningtopredict.github.io/","url_abs":"https://arxiv.org/abs/1910.13038v2","url_pdf":"https://arxiv.org/pdf/1910.13038v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"learning-to-predict-without-looking-ahead","repo_url":"https://github.com/google/brain-tokyo-workshop","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null},{"paper_slug":"learning-to-predict-without-looking-ahead","repo_url":"https://github.com/worldanon/learntopredict","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"model-based-reinforcement-learning","task_name":"Model-based Reinforcement Learning"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1910.13038","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1910.13038"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/google/brain-tokyo-workshop","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/worldanon/learntopredict","reach":null}],"summary":{"ran_honours":1,"ran_fixture":1,"unverified":1},"by_repo_kind":{"listed":{"samples":3,"ran":2,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"91277656d56a208d","entry":"encode_result_packet","repo":"worldanon/learntopredict","repo_kind":"listed","path":"gridworld/train_grid.py","file_url":"https://github.com/worldanon/learntopredict/blob/HEAD/gridworld/train_grid.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"91277656d56a208d"}},{"code_sha256_prefix":"8b59cd0f4edbb9e3","entry":"encode_solution_packets","repo":"worldanon/learntopredict","repo_kind":"listed","path":"gridworld/train_grid.py","file_url":"https://github.com/worldanon/learntopredict/blob/HEAD/gridworld/train_grid.py","link_basis":"first_harvest_node","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"RAISES","metamorphic_tier":"well_formed","behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"8b59cd0f4edbb9e3"}},{"code_sha256_prefix":"11d42d7e546811b9","entry":"decode_solution_packet","repo":"worldanon/learntopredict","repo_kind":"listed","path":"gridworld/train_grid.py","file_url":"https://github.com/worldanon/learntopredict/blob/HEAD/gridworld/train_grid.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"11d42d7e546811b9"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}