{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/despot-online-pomdp-planning-with","title":"DESPOT: Online POMDP Planning with Regularization","arxiv_id":"1609.03250","date":"2016-09-12","proceeding":"NeurIPS 2013 12","authors":["Nan Ye","Adhiraj Somani","David Hsu","Wee Sun Lee"],"abstract":"The partially observable Markov decision process (POMDP) provides a\nprincipled general framework for planning under uncertainty, but solving POMDPs\noptimally is computationally intractable, due to the \"curse of dimensionality\"\nand the \"curse of history\". To overcome these challenges, we introduce the\nDeterminized Sparse Partially Observable Tree (DESPOT), a sparse approximation\nof the standard belief tree, for online planning under uncertainty. A DESPOT\nfocuses online planning on a set of randomly sampled scenarios and compactly\ncaptures the \"execution\" of all policies under these scenarios. We show that\nthe best policy obtained from a DESPOT is near-optimal, with a regret bound\nthat depends on the representation size of the optimal policy. Leveraging this\nresult, we give an anytime online planning algorithm, which searches a DESPOT\nfor a policy that optimizes a regularized objective function. Regularization\nbalances the estimated value of a policy under the sampled scenarios and the\npolicy size, thus avoiding overfitting. The algorithm demonstrates strong\nexperimental results, compared with some of the best online POMDP algorithms\navailable. It has also been incorporated into an autonomous driving system for\nreal-time vehicle control. The source code for the algorithm is available\nonline.","url_abs":"http://arxiv.org/abs/1609.03250v3","url_pdf":"http://arxiv.org/pdf/1609.03250v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"despot-online-pomdp-planning-with","repo_url":"https://github.com/AdaCompNUS/despot","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"autonomous-driving","task_name":"Autonomous Driving"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1609.03250","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}