{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/ranked-reward-enabling-self-play","title":"Ranked Reward: Enabling Self-Play Reinforcement Learning for Combinatorial Optimization","arxiv_id":"1807.01672","date":"2018-07-04","proceeding":null,"authors":["Alexandre Laterre","Yunguan Fu","Mohamed Khalil Jabri","Alain-Sam Cohen","David Kas","Karl Hajjar","Torbjorn S. Dahl","Amine Kerkeni","Karim Beguir"],"abstract":"Adversarial self-play in two-player games has delivered impressive results\nwhen used with reinforcement learning algorithms that combine deep neural\nnetworks and tree search. Algorithms like AlphaZero and Expert Iteration learn\ntabula-rasa, producing highly informative training data on the fly. However,\nthe self-play training strategy is not directly applicable to single-player\ngames. Recently, several practically important combinatorial optimisation\nproblems, such as the travelling salesman problem and the bin packing problem,\nhave been reformulated as reinforcement learning problems, increasing the\nimportance of enabling the benefits of self-play beyond two-player games. We\npresent the Ranked Reward (R2) algorithm which accomplishes this by ranking the\nrewards obtained by a single agent over multiple games to create a relative\nperformance metric. Results from applying the R2 algorithm to instances of a\ntwo-dimensional and three-dimensional bin packing problems show that it\noutperforms generic Monte Carlo tree search, heuristic algorithms and integer\nprogramming solvers. We also present an analysis of the ranked reward\nmechanism, in particular, the effects of problem instances with varying\ndifficulty and different ranking thresholds.","url_abs":"http://arxiv.org/abs/1807.01672v3","url_pdf":"http://arxiv.org/pdf/1807.01672v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"ranked-reward-enabling-self-play","repo_url":"https://github.com/VasaKiDD/alpha-regex","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"ranked-reward-enabling-self-play","repo_url":"https://github.com/wh1992v/R2RRMopionSolitaire","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"combinatorial-optimization","task_name":"Combinatorial Optimization"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[{"method_slug":"alphazero","method_name":"AlphaZero"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1807.01672","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}