TD-Gammon
Backgammon program by Gerald Tesauro that used temporal-difference learning to reach expert level, demonstrating reinforcement learning on a complex game.
1 milestone
TD-Gammon Developed by Gerald Tesauro at IBM
In 1992, Gerald Tesauro at IBM Thomas J. Watson Research Center developed TD-Gammon, a backgammon program that trained itself through self-play using temporal-difference learning applied to a multilayer neural network, reaching a standard of play close to that of strong human experts.