{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/nonparametric-stochastic-compositional","title":"Nonparametric Stochastic Compositional Gradient Descent for Q-Learning in Continuous Markov Decision Problems","arxiv_id":"1804.07323","date":"2018-04-19","proceeding":null,"authors":["Alec Koppel","Ekaterina Tolstaya","Ethan Stump","Alejandro Ribeiro"],"abstract":"We consider Markov Decision Problems defined over continuous state and action\nspaces, where an autonomous agent seeks to learn a map from its states to\nactions so as to maximize its long-term discounted accumulation of rewards. We\naddress this problem by considering Bellman's optimality equation defined over\naction-value functions, which we reformulate into a nested non-convex\nstochastic optimization problem defined over a Reproducing Kernel Hilbert Space\n(RKHS). We develop a functional generalization of stochastic quasi-gradient\nmethod to solve it, which, owing to the structure of the RKHS, admits a\nparameterization in terms of scalar weights and past state-action pairs which\ngrows proportionately with the algorithm iteration index. To ameliorate this\ncomplexity explosion, we apply Kernel Orthogonal Matching Pursuit to the\nsequence of kernel weights and dictionaries, which yields a controllable error\nin the descent direction of the underlying optimization method. We prove that\nthe resulting algorithm, called KQ-Learning, converges with probability 1 to a\nstationary point of this problem, yielding a fixed point of the Bellman\noptimality operator under the hypothesis that it belongs to the RKHS. Under\nconstant learning rates, we further obtain convergence to a small Bellman error\nthat depends on the chosen learning rates. Numerical evaluation on the\nContinuous Mountain Car and Inverted Pendulum tasks yields convergent\nparsimonious learned action-value functions, policies that are competitive with\nthe state of the art, and exhibit reliable, reproducible learning behavior.","url_abs":"http://arxiv.org/abs/1804.07323v1","url_pdf":"http://arxiv.org/pdf/1804.07323v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"nonparametric-stochastic-compositional","repo_url":"https://github.com/katetolstaya/kernelrl","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"q-learning","task_name":"Q-Learning"},{"task_slug":"stochastic-optimization","task_name":"Stochastic Optimization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}