{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/optimal-and-adaptive-off-policy-evaluation-in","title":"Optimal and Adaptive Off-policy Evaluation in Contextual Bandits","arxiv_id":"1612.01205","date":"2016-12-04","proceeding":"ICML 2017 8","authors":["Yu-Xiang Wang","Alekh Agarwal","Miroslav Dudik"],"abstract":"We study the off-policy evaluation problem---estimating the value of a target\npolicy using data collected by another policy---under the contextual bandit\nmodel. We consider the general (agnostic) setting without access to a\nconsistent model of rewards and establish a minimax lower bound on the mean\nsquared error (MSE). The bound is matched up to constants by the inverse\npropensity scoring (IPS) and doubly robust (DR) estimators. This highlights the\ndifficulty of the agnostic contextual setting, in contrast with multi-armed\nbandits and contextual bandits with access to a consistent reward model, where\nIPS is suboptimal. We then propose the SWITCH estimator, which can use an\nexisting reward model (not necessarily consistent) to achieve a better\nbias-variance tradeoff than IPS and DR. We prove an upper bound on its MSE and\ndemonstrate its benefits empirically on a diverse collection of data sets,\noften outperforming prior work by orders of magnitude.","url_abs":"http://arxiv.org/abs/1612.01205v2","url_pdf":"http://arxiv.org/pdf/1612.01205v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"optimal-and-adaptive-off-policy-evaluation-in","repo_url":"https://github.com/facebookresearch/Horizon","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}},{"paper_slug":"optimal-and-adaptive-off-policy-evaluation-in","repo_url":"https://github.com/facebookresearch/ReAgent","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}},{"paper_slug":"optimal-and-adaptive-off-policy-evaluation-in","repo_url":"https://github.com/PlaytikaOSS/pybandits","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"none","reach":{"status":"ok","spdx":"NOASSERTION"}}],"tasks":[{"task_slug":"multi-armed-bandits","task_name":"Multi-Armed Bandits"},{"task_slug":"off-policy-evaluation","task_name":"Off-policy evaluation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1612.01205","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}