{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/multi-step-off-policy-learning-without","title":"Multi-step Off-policy Learning Without Importance Sampling Ratios","arxiv_id":"1702.03006","date":"2017-02-09","proceeding":null,"authors":["Ashique Rupam Mahmood","Huizhen Yu","Richard S. Sutton"],"abstract":"To estimate the value functions of policies from exploratory data, most\nmodel-free off-policy algorithms rely on importance sampling, where the use of\nimportance sampling ratios often leads to estimates with severe variance. It is\nthus desirable to learn off-policy without using the ratios. However, such an\nalgorithm does not exist for multi-step learning with function approximation.\nIn this paper, we introduce the first such algorithm based on\ntemporal-difference (TD) learning updates. We show that an explicit use of\nimportance sampling ratios can be eliminated by varying the amount of\nbootstrapping in TD updates in an action-dependent manner. Our new algorithm\nachieves stability using a two-timescale gradient-based TD update. A prior\nalgorithm based on lookup table representation called Tree Backup can also be\nretrieved using action-dependent bootstrapping, becoming a special case of our\nalgorithm. In two challenging off-policy tasks, we demonstrate that our\nalgorithm is stable, effectively avoids the large variance issue, and can\nperform substantially better than its state-of-the-art counterpart.","url_abs":"http://arxiv.org/abs/1702.03006v1","url_pdf":"http://arxiv.org/pdf/1702.03006v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"multi-step-off-policy-learning-without","repo_url":"https://github.com/sinaghiassian/OffpolicyAlgorithms","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1702.03006","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}