{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/making-deep-q-learning-methods-robust-to-time","title":"Making Deep Q-learning methods robust to time discretization","arxiv_id":"1901.09732","date":"2019-01-28","proceeding":null,"authors":["Corentin Tallec","Léonard Blier","Yann Ollivier"],"abstract":"Despite remarkable successes, Deep Reinforcement Learning (DRL) is not robust\nto hyperparameterization, implementation details, or small environment changes\n(Henderson et al. 2017, Zhang et al. 2018). Overcoming such sensitivity is key\nto making DRL applicable to real world problems. In this paper, we identify\nsensitivity to time discretization in near continuous-time environments as a\ncritical factor; this covers, e.g., changing the number of frames per second,\nor the action frequency of the controller. Empirically, we find that\nQ-learning-based approaches such as Deep Q- learning (Mnih et al., 2015) and\nDeep Deterministic Policy Gradient (Lillicrap et al., 2015) collapse with small\ntime steps. Formally, we prove that Q-learning does not exist in continuous\ntime. We detail a principled way to build an off-policy RL algorithm that\nyields similar performances over a wide range of time discretizations, and\nconfirm this robustness empirically.","url_abs":"http://arxiv.org/abs/1901.09732v2","url_pdf":"http://arxiv.org/pdf/1901.09732v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"making-deep-q-learning-methods-robust-to-time","repo_url":"https://github.com/ctallec/continuous-rl","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"deep-reinforcement-learning","task_name":"Deep Reinforcement Learning"},{"task_slug":"q-learning","task_name":"Q-Learning"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"sensitivity","task_name":"Sensitivity"}],"methods":[{"method_slug":"q-learning","method_name":"Q-Learning"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1901.09732","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1901.09732"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/ctallec/continuous-rl","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"unverified":1},"by_repo_kind":{"listed":{"samples":1,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"a8989b4bdc4d5c51","entry":"setup_optimizer","repo":"ctallec/continuous-rl","repo_kind":"listed","path":"code/optimizer.py","file_url":"https://github.com/ctallec/continuous-rl/blob/HEAD/code/optimizer.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"a8989b4bdc4d5c51"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}