Benchmark or contest
A decisive result against a recognised benchmark or human opponent.
12 milestones.
DeepMind's AlphaFold 2 Achieves Highest-Accuracy Results at CASP14 Protein Structure Prediction Competition
In November–December 2020, DeepMind's AlphaFold 2 system achieved a median Global Distance Test score of approximately 92.4 across all CASP14 targets, far surpassing the next-best group, in a result that computational biologists described as largely solving the 50-year-old protein-folding problem for single-chain proteins.
Google AI Model Matches or Exceeds Radiologist Performance in Lung Cancer Detection from CT Scans
In May 2019, researchers at Google Health and Northwestern Medicine published a deep-learning model in Nature Medicine that detected malignant lung nodules in low-dose CT scans, matching or exceeding the performance of six radiologists on a held-out dataset, with fewer false positives and false negatives when prior scans were unavailable.
OpenAI Released OpenAI Gym, a Toolkit for Reinforcement Learning Research
In April 2016, OpenAI publicly released OpenAI Gym, an open-source toolkit providing a standardised collection of environments for developing and benchmarking reinforcement learning algorithms, lowering the barrier to reproducible RL research.
DeepMind Publishes AlphaGo, a Deep Reinforcement Learning System That Defeated Professional Go Players
In January 2016, researchers at Google DeepMind published a paper in Nature describing AlphaGo, a system combining deep convolutional neural networks with Monte Carlo tree search and reinforcement learning that defeated the European Go champion Fan Hui 5–0, marking the first time a computer program had beaten a professional Go player at full-board Go.
Schaft Inc Robot Wins DARPA Robotics Challenge Trials 2013
In December 2013, Schaft Inc, a Japanese robotics company acquired by Google in October 2013, won the DARPA Robotics Challenge Trials at Homestead Miami Speedway, Florida, scoring 27 out of 32 points across eight disaster-response tasks and finishing ahead of 15 other teams.
AlexNet and Deep Convolutional Neural Networks in Large-Scale Image Classification
In September 2012, Alex Krizhevsky, Ilya Sutskever and Geoffrey Hinton at the University of Toronto submitted a paper describing AlexNet, a deep convolutional neural network that achieved a top-5 error rate of 15.3% on the ImageNet Large Scale Visual Recognition Challenge, outperforming the next-best entry by more than 10 percentage points.
IBM Watson Defeats Human Champions on Jeopardy!
In February 2011, IBM's Watson system defeated Jeopardy! champions Ken Jennings and Brad Rutter across three televised episodes, demonstrating that a computer could parse ambiguous natural-language clues and retrieve factual answers competitively against expert human players.
Netflix Prize Competition Launched by Netflix
In October 2006, Netflix launched the Netflix Prize, an open competition offering $1,000,000 USD to any team that could improve the accuracy of the company's Cinematch recommendation algorithm by at least 10% on a supplied ratings dataset, measured by root mean squared error.
Stanford Racing Team's Stanley Wins DARPA Grand Challenge 2005
On 8 October 2005, a Stanford University team led by Sebastian Thrun entered Stanley, a modified Volkswagen Touareg, in the DARPA Grand Challenge. Stanley completed the 131.6-mile (211.8 km) Mojave Desert course autonomously in under 7 hours, finishing first and winning the $2 million prize.
Deep Blue Defeats Garry Kasparov in Six-Game Rematch
In May 1997, IBM's Deep Blue chess-playing system defeated reigning world champion Garry Kasparov over a six-game match by a score of 3½–2½, becoming the first computer system to defeat a reigning world chess champion under standard tournament conditions.
Chinook Defeats Marion Tinsley to Win the American Checkers Federation and World Checkers Federation Championship
In 1994, Chinook, a checkers-playing program developed by Jonathan Schaeffer and colleagues at the University of Alberta, became world champion after Marion Tinsley withdrew from their match due to illness, making it the first computer program to win a human world championship in any board game.
INTERNIST-I Developed by Jack D. Myers and Harry E. Pople Jr.
Jack D. Myers and Harry E. Pople Jr. at the University of Pittsburgh developed INTERNIST-I, an expert system for internal medicine diagnosis, with its principal public evaluation published in the New England Journal of Medicine in 1982. The system encoded knowledge of roughly 500 diseases and 3,500 symptoms, demonstrating that algorithmic clinical reasoning could approach specialist performance on complex cases.