{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/adversarial-online-multi-task-reinforcement","title":"Adversarial Online Multi-Task Reinforcement Learning","arxiv_id":"2301.04268","date":"2023-01-11","proceeding":null,"authors":["Quan Nguyen","Nishant A. Mehta"],"abstract":"We consider the adversarial online multi-task reinforcement learning setting, where in each of $K$ episodes the learner is given an unknown task taken from a finite set of $M$ unknown finite-horizon MDP models. The learner's objective is to minimize its regret with respect to the optimal policy for each task. We assume the MDPs in $\\mathcal{M}$ are well-separated under a notion of $\\lambda$-separability, and show that this notion generalizes many task-separability notions from previous works. We prove a minimax lower bound of $\\Omega(K\\sqrt{DSAH})$ on the regret of any learning algorithm and an instance-specific lower bound of $\\Omega(\\frac{K}{\\lambda^2})$ in sample complexity for a class of uniformly-good cluster-then-learn algorithms. We use a novel construction called 2-JAO MDP for proving the instance-specific lower bound. The lower bounds are complemented with a polynomial time algorithm that obtains $\\tilde{O}(\\frac{K}{\\lambda^2})$ sample complexity guarantee for the clustering phase and $\\tilde{O}(\\sqrt{MK})$ regret guarantee for the learning phase, indicating that the dependency on $K$ and $\\frac{1}{\\lambda^2}$ is tight.","url_abs":"https://arxiv.org/abs/2301.04268v1","url_pdf":"https://arxiv.org/pdf/2301.04268v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"adversarial-online-multi-task-reinforcement","repo_url":"https://github.com/ngmq/adversarial-online-multi-task-reinforcement-learning","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}