Deep learning

37 milestones used this technique.

3D Gaussian Splatting achieves real-time novel-view synthesis at 1080p

Bernhard Kerbl and colleagues applied 3D Gaussian splatting to novel-view synthesis with end-to-end optimisation and adaptive density control, achieving high-quality results at 1080p resolution and real-time frame rates, without the slow neural rendering that previous approaches required.

FlashAttention-2 roughly doubles attention speed on A100 GPUs

Tri Dao's FlashAttention-2 improved GPU work partitioning to deliver roughly 2x the speed of FlashAttention, reaching 50–73% of the theoretical maximum FLOPs/s in the configurations benchmarked in the paper, and up to 225 TFLOPs/s per A100 in the specific GPT-style training configurations reported.

QLoRA enables finetuning of 65B-parameter models on a single 48GB GPU

Tim Dettmers and colleagues showed that a 65-billion-parameter language model could be finetuned on a single 48GB GPU without measurable quality loss on the benchmarks tested, by combining 4-bit quantisation with low-rank adapter training.

LLaMA matches leading models on public data with far fewer parameters

LLaMA, a collection of foundation language models from 7B to 65B parameters trained exclusively on publicly available data, was submitted to arXiv on 27 February 2023. The 13B model outperformed GPT-3 at 175B parameters on most benchmarks, and the weights were released to the research community.

ControlNet adds structured spatial conditioning to pretrained text-to-image diffusion models

Lvmin Zhang, Anyi Rao and Maneesh Agrawala introduced ControlNet, a neural network architecture that adds spatial conditioning controls such as edges, depth, segmentation and human pose to large pretrained text-to-image diffusion models without degrading their existing capabilities.

BLOOM: open-access 176B-parameter multilingual language model

The BigScience Workshop released BLOOM, a 176-billion-parameter language model trained across 46 natural and 13 programming languages, made freely available under the Responsible AI License at a scale that had only recently begun to become accessible, and had not previously been available with multilingual training data.

Flan-PaLM: instruction finetuning scaled across tasks, model sizes and families

Researchers showed that finetuning large language models on instruction-phrased datasets improves performance across benchmarks, with Flan-PaLM 540B trained on 1.8K tasks scoring 75.2% on five-shot MMLU and outperforming its base model by 9.4% on average.

Flamingo: few-shot visual language model for interleaved images, video and text

Researchers introduced Flamingo, a family of Visual Language Models that could handle interleaved images, video and text, achieving state-of-the-art few-shot performance on many benchmarks without task-specific fine-tuning.

InstructGPT: aligning language models with human feedback at scale

Researchers showed that fine-tuning GPT-3 with human feedback produced a 1.3B parameter model whose outputs labellers preferred over those of the 175B GPT-3, pointing toward a practical method for aligning language models more closely with expressed human preferences.

Multiresolution hash encoding cuts neural graphics training to seconds on a single GPU

Thomas Müller, Alex Evans, Christoph Schied, and Alexander Keller introduced a multiresolution hash encoding that trains neural graphics primitives in seconds and renders at 1920×1080 in tens of milliseconds, achieving a combined speedup of several orders of magnitude when the encoding and optimised CUDA kernels are used together.

BrainGate researchers decode imagined handwriting from neural signals to enable high-speed text communication

In May 2021, Francis R. Willett and colleagues at Stanford University and Howard Hughes Medical Institute published results showing that a BrainGate2 intracortical electrode array could decode imagined handwriting movements in a paralysed person at 90 characters per minute with 94.1% raw accuracy, substantially exceeding prior neural-interface typing rates.

DeepMind's AlphaFold 2 Achieves Highest-Accuracy Results at CASP14 Protein Structure Prediction Competition

In November–December 2020, DeepMind's AlphaFold 2 system achieved a median Global Distance Test score of approximately 92.4 across all CASP14 targets, far surpassing the next-best group, in a result that computational biologists described as largely solving the 50-year-old protein-folding problem for single-chain proteins.

Waymo One Launches Fully Driverless Rides to the General Public in Phoenix, Arizona

On 8 October 2020, Waymo opened its Waymo One ride-hailing service to the general public in the greater Phoenix, Arizona area, operating without a safety driver in the vehicle, the first time a commercial autonomous vehicle service had done so at public scale.

Once-for-All: Train One Network and Specialize It for Efficient Deployment

Han Cai, Chuang Gan, Tianhao Chen, and Song Han at MIT published Once-for-All at ICLR 2020, presenting a method to train a single neural network once and then derive specialised sub-networks for diverse hardware platforms without retraining, reducing the computational cost of neural architecture search by orders of magnitude.

MIT Researchers Use Machine Learning to Identify Halicin, an Antibiotic Effective Against Drug-Resistant Bacteria

On 20 February 2020, James Collins and colleagues at MIT published research in Cell describing a deep-learning model trained to predict antibiotic activity; the model identified halicin, a compound previously investigated for diabetes treatment, as a potent broad-spectrum antibiotic capable of killing several drug-resistant bacterial strains.

Microsoft Research released Turing Natural Language Generation (T-NLG), a 17-billion-parameter language model

In February 2020, Microsoft Research announced Turing Natural Language Generation (T-NLG), a 17-billion-parameter autoregressive language model trained using the Megatron-LM framework. At the time of release it was the largest publicly disclosed language model and achieved state-of-the-art results on question-answering and summarisation benchmarks.

Facebook AI Research Publishes GrokNet, a Unified Computer Vision Model for Commerce Understanding

In 2020, researchers at Facebook AI published GrokNet, a unified deep learning system for product understanding in commerce settings, capable of recognising object categories, attributes such as colour and material, and brand information from product images at scale across Facebook Shops.

Facebook AI Research Releases Detectron2

In October 2019, Facebook AI Research released Detectron2, an open-source object detection and segmentation framework built on PyTorch, supporting algorithms including Mask R-CNN, DensePose, and panoptic feature pyramid networks, replacing the earlier Caffe2-based Detectron.

Google AI Model Matches or Exceeds Radiologist Performance in Lung Cancer Detection from CT Scans

In May 2019, researchers at Google Health and Northwestern Medicine published a deep-learning model in Nature Medicine that detected malignant lung nodules in low-dose CT scans, matching or exceeding the performance of six radiologists on a held-out dataset, with fewer false positives and false negatives when prior scans were unavailable.

LOVOT Companion Robot Unveiled by Groove X

In December 2018, Groove X, a Japanese robotics company founded by Kaname Hayashi, unveiled LOVOT, a companion robot designed to elicit emotional attachment rather than perform practical tasks, equipped with more than 50 sensors, a thermal camera array, and a neural-processing unit to recognise and respond to human behaviour.

Waymo One Commercial Ride-Hailing Service Launch

In December 2018, Waymo LLC launched Waymo One, a fare-charging autonomous ride-hailing service operating in the Greater Phoenix, Arizona area, marking the first time a driverless vehicle service had been made available to paying members of the public in the United States.

BERT introduces masked bidirectional pre-training for language understanding

Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova introduced BERT, a language model that pre-trains on both left and right context simultaneously, achieving new best results on eleven natural language processing tasks.

Facebook Deploys AI-Assisted Suicide Prevention Detection Across Live Video

In November 2017, Facebook announced the global expansion of an AI system designed to detect signs of suicidal intent in posts and Live videos, using pattern recognition trained on reports flagged by human reviewers to surface at-risk content to its Community Operations team and connect users with crisis resources.

Facebook AI Research Published StarSpace: Embed All The Things!

In September 2017, Ledell Wu and colleagues at Facebook AI Research published StarSpace (arXiv:1709.03856), a general-purpose neural embedding model capable of learning entity representations across tasks including text classification, ranking, and collaborative filtering, without task-specific architecture changes.

Movidius (Intel) Launches Neural Compute Stick

In July 2017, Movidius, an Intel subsidiary, released the Movidius Neural Compute Stick, a USB-form-factor device housing the Myriad 2 Vision Processing Unit, enabling developers to run inference from trained deep neural networks on low-power edge hardware without a remote server.

Caffe2Go: Facebook's On-Device Neural Style Transfer for Mobile Video

In November 2016, researchers at Facebook AI Research published Caffe2Go, a compressed deep-learning framework that ran neural style-transfer models entirely on iOS and Android devices without sending video frames to a server, enabling real-time artistic video effects on mobile hardware.

WaveNet: A Generative Model for Raw Audio, by DeepMind

In September 2016, researchers at Google DeepMind published WaveNet, a deep generative model that synthesises raw audio waveforms sample-by-sample using dilated causal convolutions. In evaluations on English and Mandarin speech, WaveNet reduced the gap between human speech and machine synthesis by more than 50 per cent compared with the best previous text-to-speech systems.

DeepMind Publishes AlphaGo, a Deep Reinforcement Learning System That Defeated Professional Go Players

In January 2016, researchers at Google DeepMind published a paper in Nature describing AlphaGo, a system combining deep convolutional neural networks with Monte Carlo tree search and reinforcement learning that defeated the European Go champion Fan Hui 5–0, marking the first time a computer program had beaten a professional Go player at full-board Go.

Deep residual networks make training at 100+ layers practical with identity shortcuts

Kaiming He and colleagues introduced residual learning, letting networks train at depths of up to 152 layers. Their ResNet won first place on five tracks at the ILSVRC and COCO 2015 competitions, including ImageNet classification, detection, localisation, and COCO detection and segmentation.

DeepMind's DQN learns to play Atari games from raw pixels

DeepMind's Deep Q-Networks algorithm learned to play Atari 2600 games directly from raw pixels, matching or exceeding the score of a human tester on roughly half of the games tested, without any prior knowledge of the rules.

Kingma and Ba present the Adam optimiser

Diederik P. Kingma and Jimmy Ba presented Adam, an optimisation algorithm combining and extending existing adaptive gradient methods, adapting learning rates using bias-corrected estimates of the first and second moments of gradients, making it efficient and broadly practical, particularly where heavy hyperparameter tuning is not feasible.

Bahdanau, Cho and Bengio introduce soft attention for neural machine translation

Dzmitry Bahdanau, Kyunghyun Cho and Yoshua Bengio proposed an attention mechanism letting neural translation models search source sentences dynamically, rather than compressing everything into a single fixed-length vector, achieving performance comparable to phrase-based systems on English-to-French translation.

Goodfellow and colleagues propose generative adversarial networks

Ian Goodfellow and seven co-authors proposed training two neural networks against each other: one generating samples, one judging them. The setup, requiring only backpropagation and no Markov chains, could recover the training data distribution under idealised theoretical assumptions.

AlexNet and Deep Convolutional Neural Networks in Large-Scale Image Classification

In September 2012, Alex Krizhevsky, Ilya Sutskever and Geoffrey Hinton at the University of Toronto submitted a paper describing AlexNet, a deep convolutional neural network that achieved a top-5 error rate of 15.3% on the ImageNet Large Scale Visual Recognition Challenge, outperforming the next-best entry by more than 10 percentage points.

Dropout introduced as a regularisation method for neural networks

Srivastava, Hinton, Krizhevsky, Sutskever and Salakhutdinov introduced dropout, a method that randomly omits units during training to stop neural networks from overfitting. It set new records on several specific speech and object recognition benchmarks at the time of publication.

Google Brain Founded by Andrew Ng and Jeff Dean

In 2011, Andrew Ng and Jeff Dean co-founded Google Brain, an internal research group at Google dedicated to large-scale deep learning. The project demonstrated that deep neural networks trained on substantial compute could learn useful representations without labelled data, reshaping how the industry approached machine learning research.

ImageNet: a large-scale hierarchical image database for object recognition

Researchers led by Li Fei-Fei introduced ImageNet, a dataset organised around more than 100,000 concept categories and aimed at providing roughly 1,000 human-annotated images per category, on which the ILSVRC challenge was later built.