Self-play
Training technique where an agent learns by playing against itself, generating its own experience data without requiring human opponents or labeled examples.
1 milestone
TD-Gammon Developed by Gerald Tesauro at IBM
In 1992, Gerald Tesauro at IBM Thomas J. Watson Research Center developed TD-Gammon, a backgammon program that trained itself through self-play using temporal-difference learning applied to a multilayer neural network, reaching a standard of play close to that of strong human experts.