{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/model-free-adaptive-optimal-control-of","title":"Model-Free Adaptive Optimal Control of Episodic Fixed-Horizon Manufacturing Processes using Reinforcement Learning","arxiv_id":"1809.06646","date":"2018-09-18","proceeding":null,"authors":["Johannes Dornheim","Norbert Link","Peter Gumbsch"],"abstract":"A self-learning optimal control algorithm for episodic fixed-horizon\nmanufacturing processes with time-discrete control actions is proposed and\nevaluated on a simulated deep drawing process. The control model is built\nduring consecutive process executions under optimal control via reinforcement\nlearning, using the measured product quality as reward after each process\nexecution. Prior model formulation, which is required by state-of-the-art\nalgorithms from model predictive control and approximate dynamic programming,\nis therefore obsolete. This avoids several difficulties namely in system\nidentification, accurate modelling, and runtime complexity, that arise when\ndealing with processes subject to nonlinear dynamics and stochastic influences.\nInstead of using pre-created process and observation models, value\nfunction-based reinforcement learning algorithms build functions of expected\nfuture reward, which are used to derive optimal process control decisions. The\nexpectation functions are learned online, by interacting with the process. The\nproposed algorithm takes stochastic variations of the process conditions into\naccount and is able to cope with partial observability. A Q-learning-based\nmethod for adaptive optimal control of partially observable episodic\nfixed-horizon manufacturing processes is developed and studied. The resulting\nalgorithm is instantiated and evaluated by applying it to a simulated\nstochastic optimal control problem in metal sheet deep drawing.","url_abs":"http://arxiv.org/abs/1809.06646v3","url_pdf":"http://arxiv.org/pdf/1809.06646v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"model-free-adaptive-optimal-control-of","repo_url":"https://github.com/johannes-dornheim/Reinforce-FE","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"model-predictive-control","task_name":"Model Predictive Control"},{"task_slug":"q-learning","task_name":"Q-Learning"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"self-learning","task_name":"Self-Learning"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}