{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/reinforcement-learning-with-latent-flow-1","title":"Reinforcement Learning with Latent Flow","arxiv_id":"2101.01857","date":"2021-01-06","proceeding":"NeurIPS 2021 12","authors":["Wenling Shang","Xiaofei Wang","Aravind Srinivas","Aravind Rajeswaran","Yang Gao","Pieter Abbeel","Michael Laskin"],"abstract":"Temporal information is essential to learning effective policies with Reinforcement Learning (RL). However, current state-of-the-art RL algorithms either assume that such information is given as part of the state space or, when learning from pixels, use the simple heuristic of frame-stacking to implicitly capture temporal information present in the image observations. This heuristic is in contrast to the current paradigm in video classification architectures, which utilize explicit encodings of temporal information through methods such as optical flow and two-stream architectures to achieve state-of-the-art performance. Inspired by leading video classification architectures, we introduce the Flow of Latents for Reinforcement Learning (Flare), a network architecture for RL that explicitly encodes temporal information through latent vector differences. We show that Flare (i) recovers optimal performance in state-based RL without explicit access to the state velocity, solely with positional state information, (ii) achieves state-of-the-art performance on pixel-based challenging continuous control tasks within the DeepMind control benchmark suite, namely quadruped walk, hopper hop, finger turn hard, pendulum swing, and walker run, and is the most sample efficient model-free pixel-based RL algorithm, outperforming the prior model-free state-of-the-art by 1.9X and 1.5X on the 500k and 1M step benchmarks, respectively, and (iv), when augmented over rainbow DQN, outperforms this state-of-the-art level baseline on 5 of 8 challenging Atari games at 100M time step benchmark.","url_abs":"https://arxiv.org/abs/2101.01857v1","url_pdf":"https://arxiv.org/pdf/2101.01857v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"reinforcement-learning-with-latent-flow-1","repo_url":"https://github.com/WendyShang/flare","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":null},{"paper_slug":"reinforcement-learning-with-latent-flow-1","repo_url":"https://github.com/WendyShang/dqn_zoo","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"jax","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"atari-games","task_name":"Atari Games"},{"task_slug":"continuous-control","task_name":"Continuous Control"},{"task_slug":"montezumas-revenge","task_name":"Montezuma's Revenge"},{"task_slug":"optical-flow-estimation","task_name":"Optical Flow Estimation"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"video-classification","task_name":"Video Classification"},{"task_slug":"continuous-control","task_name":"continuous-control"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[{"method_slug":"convolution","method_name":"Convolution"},{"method_slug":"dqn","method_name":"DQN"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"q-learning","method_name":"Q-Learning"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/montezuma-s-revenge-on-atari-2600-montezuma-s","task":"Montezuma's Revenge","dataset":"Atari 2600 Montezuma's Revenge","model":"Flare","rank_in_archive_order":1,"of":3,"metrics":{"Average Return (NoOp)":"1668"},"uses_additional_data":false},{"leaderboard":"/sota/montezuma-s-revenge-on-atari-2600-montezuma-s","task":"Montezuma's Revenge","dataset":"Atari 2600 Montezuma's Revenge","model":"Rainbow (tuned)","rank_in_archive_order":2,"of":3,"metrics":{"Average Return (NoOp)":"900"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2101.01857","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2101.01857"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/WendyShang/dqn_zoo","reach":{"status":"unanswered"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/WendyShang/flare","reach":null}],"summary":{"ran_fixture":1},"by_repo_kind":{"official":{"samples":1,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"9417c152a8c940ae","entry":"center_crop_images","repo":"WendyShang/flare","repo_kind":"official","path":"train_ddp.py","file_url":"https://github.com/WendyShang/flare/blob/HEAD/train_ddp.py","link_basis":"first_harvest_node","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"RAISES","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"9417c152a8c940ae"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}