{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/td-or-not-td-analyzing-the-role-of-temporal","title":"TD or not TD: Analyzing the Role of Temporal Differencing in Deep Reinforcement Learning","arxiv_id":"1806.01175","date":"2018-06-04","proceeding":"ICLR 2018 1","authors":["Artemij Amiranashvili","Alexey Dosovitskiy","Vladlen Koltun","Thomas Brox"],"abstract":"Our understanding of reinforcement learning (RL) has been shaped by\ntheoretical and empirical results that were obtained decades ago using tabular\nrepresentations and linear function approximators. These results suggest that\nRL methods that use temporal differencing (TD) are superior to direct Monte\nCarlo estimation (MC). How do these results hold up in deep RL, which deals\nwith perceptually complex environments and deep nonlinear models? In this\npaper, we re-examine the role of TD in modern deep RL, using specially designed\nenvironments that control for specific factors that affect performance, such as\nreward sparsity, reward delay, and the perceptual complexity of the task. When\ncomparing TD with infinite-horizon MC, we are able to reproduce classic results\nin modern settings. Yet we also find that finite-horizon MC is not inferior to\nTD, even when rewards are sparse or delayed. This makes MC a viable alternative\nto TD in deep RL.","url_abs":"http://arxiv.org/abs/1806.01175v1","url_pdf":"http://arxiv.org/pdf/1806.01175v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"td-or-not-td-analyzing-the-role-of-temporal","repo_url":"https://github.com/lmb-freiburg/td-or-not-td","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"tf","reach":null}],"tasks":[{"task_slug":"deep-reinforcement-learning","task_name":"Deep Reinforcement Learning"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1806.01175","atlas_url":"https://app.syntology.ai/?focus=1806.01175","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}