Papers › Deep Reinforcement Learning from Self-Play in Imperfect-Information Games

Deep Reinforcement Learning from Self-Play in Imperfect-Information Games

3 Mar 2016arXiv:1603.01121archive 2025-07-28

Johannes Heinrich, David Silver

Many real-world applications can be described as large-scale games of imperfect information. To deal with these challenging domains, prior work has focused on computing Nash equilibria in a handcrafted abstraction of the domain. In this paper we introduce the first scalable end-to-end approach to learning approximate Nash equilibria without prior domain knowledge. Our method combines fictitious self-play with deep reinforcement learning. When applied to Leduc poker, Neural Fictitious Self-Play (NFSP) approached a Nash equilibrium, whereas common reinforcement learning methods diverged. In Limit Texas Holdem, a poker game of real-world scale, NFSP learnt a strategy that approached the performance of state-of-the-art, superhuman algorithms based on significant domain expertise.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

EricSteinberger/DREAM mentioned on GitHubMIT report
IAARhub/TrucoAnalytics mentioned on GitHub report
TinkeringCode/Neural-Fictitous-Self-Play mentioned on GitHubpytorchMIT report
heidekrueger/bnelearn mentioned on GitHubpytorchGPL-3.0 report
jsanderink/tue mentioned on GitHubpytorchMIT report
quantumiracle/mars mentioned on GitHubpytorchApache-2.0 report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Card GamesDeep Reinforcement LearningGame of PokerReinforcement LearningReinforcement Learning (RL)reinforcement-learning

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections