{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/sai-a-sensible-artificial-intelligence-that","title":"SAI, a Sensible Artificial Intelligence that plays Go","arxiv_id":"1809.03928","date":"2018-09-11","proceeding":null,"authors":["Francesco Morandin","Gianluca Amato","Rosa Gini","Carlo Metta","Maurizio Parton","Gian-Carlo Pascutto"],"abstract":"We propose a multiple-komi modification of the AlphaGo Zero/Leela Zero\nparadigm. The winrate as a function of the komi is modeled with a\ntwo-parameters sigmoid function, so that the neural network must predict just\none more variable to assess the winrate for all komi values. A second novel\nfeature is that training is based on self-play games that occasionally branch\n-- with changed komi -- when the position is uneven. With this setting,\nreinforcement learning is showed to work on 7x7 Go, obtaining very strong\nplaying agents. As a useful byproduct, the sigmoid parameters given by the\nnetwork allow to estimate the score difference on the board, and to evaluate\nhow much the game is decided.","url_abs":"http://arxiv.org/abs/1809.03928v2","url_pdf":"http://arxiv.org/pdf/1809.03928v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"sai-a-sensible-artificial-intelligence-that","repo_url":"https://github.com/sai-dev/sai","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":null,"task_name":"Position"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}