AI research

Learning methods

AI milestones in learning methods, part of ai research.

10 milestones

Direct Preference Optimisation reduces need for RL in LM alignment

Rafael Rafailov and colleagues introduced Direct Preference Optimization (DPO), a method that aligns language models with human preferences using only a simple classification loss, bypassing the complex reinforcement learning pipeline that existing approaches required.

QLoRA enables finetuning of 65B-parameter models on a single 48GB GPU

Tim Dettmers and colleagues showed that a 65-billion-parameter language model could be finetuned on a single 48GB GPU without measurable quality loss on the benchmarks tested, by combining 4-bit quantisation with low-rank adapter training.

DeepMind's DQN learns to play Atari games from raw pixels

DeepMind's Deep Q-Networks algorithm learned to play Atari 2600 games directly from raw pixels, matching or exceeding the score of a human tester on roughly half of the games tested, without any prior knowledge of the rules.

Kingma and Ba present the Adam optimiser

Diederik P. Kingma and Jimmy Ba presented Adam, an optimisation algorithm combining and extending existing adaptive gradient methods, adapting learning rates using bias-corrected estimates of the first and second moments of gradients, making it efficient and broadly practical, particularly where heavy hyperparameter tuning is not feasible.

Goodfellow and colleagues propose generative adversarial networks

Ian Goodfellow and seven co-authors proposed training two neural networks against each other: one generating samples, one judging them. The setup, requiring only backpropagation and no Markov chains, could recover the training data distribution under idealised theoretical assumptions.

Xinlei Chen, Abhinav Shrivastava and Abhinav Gupta at Carnegie Mellon University present NEIL (Never-Ending Image Learner) at ICCV 2013

In December 2013, Xinlei Chen, Abhinav Shrivastava and Abhinav Gupta at Carnegie Mellon University presented NEIL (Never-Ending Image Learner) at ICCV 2013, a continuously running system that autonomously mined semantic relationships between visual concepts from unlabelled web images without human supervision.

Dropout introduced as a regularisation method for neural networks

Srivastava, Hinton, Krizhevsky, Sutskever and Salakhutdinov introduced dropout, a method that randomly omits units during training to stop neural networks from overfitting. It set new records on several specific speech and object recognition benchmarks at the time of publication.

Google Brain Unsupervised Neural Network Learns to Detect Cats from YouTube Frames

In June 2012, Quoc V. Le and colleagues at Google Brain published research showing that a 1,000-machine, 16,000-core neural network trained without labels on 10 million YouTube thumbnail images spontaneously developed a neuron selectively responsive to human and cat faces, demonstrating large-scale unsupervised feature learning from unlabelled video data.

IBM TJ Watson Research Center Publishes Statistical Approach to Machine Translation

In August 1988, researchers at IBM Thomas J. Watson Research Center, including Peter F. Brown, John Cocke, Stephen A. Della Pietra, Vincent J. Della Pietra, Fredrick Jelinek, Robert L. Mercer, and Paul S. Roossin, presented a statistical framework for machine translation at COLING 1988, replacing rule-based linguistics with probabilistic models trained on bilingual text corpora.

Backpropagation Described by Arthur E. Bryson Jr. and Yu-Chi Ho

In 1969, Arthur E. Bryson Jr. and Yu-Chi Ho of Harvard University described a gradient-based optimisation procedure for multi-stage dynamic systems in their textbook Applied Optimal Control, presenting what is now recognised as an early statement of the backpropagation principle in a supervised-learning context.