ACHIEVEMENTS.AI

AI research

Architectures and models

AI milestones in architectures and models, part of ai research.

20 milestones

OpenAI released GPT-3 via private beta API

In May–June 2020, OpenAI published the GPT-3 language model in a paper by Tom B. Brown and colleagues, and began distributing private beta API access. GPT-3's 175 billion parameters made it substantially larger than any publicly described language model at the time, enabling strong few-shot performance across diverse language tasks.

Once-for-All: Train One Network and Specialize It for Efficient Deployment

Han Cai, Chuang Gan, Tianhao Chen, and Song Han at MIT published Once-for-All at ICLR 2020, presenting a method to train a single neural network once and then derive specialised sub-networks for diverse hardware platforms without retraining, reducing the computational cost of neural architecture search by orders of magnitude.

OpenAI Releases GPT-1: Improving Language Understanding by Generative Pre-Training

In June 2018, Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever at OpenAI published 'Improving Language Understanding by Generative Pre-Training', introducing GPT-1, a 117-million-parameter Transformer pretrained on BooksCorpus via unsupervised language modelling and fine-tuned on downstream tasks, outperforming task-specific models on several NLP benchmarks.

Facebook AI Research Published StarSpace: Embed All The Things!

In September 2017, Ledell Wu and colleagues at Facebook AI Research published StarSpace (arXiv:1709.03856), a general-purpose neural embedding model capable of learning entity representations across tasks including text classification, ranking, and collaborative filtering, without task-specific architecture changes.

WaveNet: A Generative Model for Raw Audio, by DeepMind

In September 2016, researchers at Google DeepMind published WaveNet, a deep generative model that synthesises raw audio waveforms sample-by-sample using dilated causal convolutions. In evaluations on English and Mandarin speech, WaveNet reduced the gap between human speech and machine synthesis by more than 50 per cent compared with the best previous text-to-speech systems.

Neural Turing Machine Introduced by Alex Graves, Greg Wayne, and Ivo Danihelka

In October 2014, Alex Graves, Greg Wayne, and Ivo Danihelka at Google DeepMind published 'Neural Turing Machines', a preprint proposing a neural network architecture augmented with an external memory matrix and differentiable read/write operations, enabling the system to learn algorithms such as sorting and copying from examples alone.

AlexNet and Deep Convolutional Neural Networks in Large-Scale Image Classification

In September 2012, Alex Krizhevsky, Ilya Sutskever and Geoffrey Hinton at the University of Toronto submitted a paper describing AlexNet, a deep convolutional neural network that achieved a top-5 error rate of 15.3% on the ImageNet Large Scale Visual Recognition Challenge, outperforming the next-best entry by more than 10 percentage points.

OpenCog Artificial General Intelligence Framework Introduced

In 2008, Ben Goertzel and colleagues at the Singularity Institute for Artificial Intelligence publicly introduced OpenCog, an open-source software framework designed to support research into artificial general intelligence by integrating multiple cognitive subsystems within a shared knowledge store called the AtomSpace.

Numenta Founded by Jeff Hawkins and Donna Dubinsky

In 2005, Jeff Hawkins and Donna Dubinsky co-founded Numenta, a research company dedicated to developing machine intelligence systems modelled on the structural and algorithmic principles of the mammalian neocortex, building on Hawkins's theoretical framework published in his 2004 book On Intelligence.

Bag of Words Applied to Computer Vision (Visual Vocabulary / Bag of Visual Words)

Josef Sivic and Andrew Zisserman at the University of Oxford applied the Bag of Words text-retrieval model to visual features in their 2003 ICCV paper 'Video Google', representing image regions as a vocabulary of visual words to enable efficient object retrieval from video.

A Neural Probabilistic Language Model by Yoshua Bengio and Colleagues

In 2003, Yoshua Bengio, Réjean Ducharme, Pascal Vincent, and Christian Janvin at the Université de Montréal published 'A Neural Probabilistic Language Model' in JMLR, demonstrating that a feed-forward neural network trained on word sequences could learn distributed word representations and outperform n-gram models on perplexity benchmarks.

Structured Light for Robust Correspondence in Active Stereo Vision

In 2002, Li Zhang, Brian Curless, and Steven M. Seitz at the University of Washington presented a method using structured light patterns projected onto scenes to establish robust stereo correspondences, enabling reliable 3D reconstruction under conditions where passive stereo fails.

Long Short-Term Memory Introduced by Sepp Hochreiter and Jürgen Schmidhuber

In 1997, Sepp Hochreiter at Technische Universität München and Jürgen Schmidhuber at IDSIA published 'Long Short-Term Memory' in Neural Computation, introducing a recurrent neural network architecture with gated memory cells that could learn dependencies across long sequences without suffering from the vanishing gradient problem.

ALICE Chatbot Created by Richard S. Wallace

In 1995, Richard S. Wallace, an independent AI researcher, created ALICE (Artificial Linguistic Internet Computer Entity), a natural-language chatbot that used a pattern-matching markup language called AIML to generate contextually plausible conversational responses, later influencing a generation of open-source chatbot development.

NETtalk Neural Network Developed by Terrence J. Sejnowski and Charles Rosenberg

Terrence J. Sejnowski of the Salk Institute and Charles Rosenberg of Princeton University developed NETtalk, a feedforward neural network trained to convert English text to speech, publishing the principal account in Complex Systems in 1987. The network learned pronunciation from examples alone, demonstrating that a multi-layer perceptron could acquire a complex linguistic skill without hand-coded rules.

SOAR Cognitive Architecture: Doctoral Dissertations by John E. Laird and Paul S. Rosenbloom, Supervised by Allen Newell

In 1983, John E. Laird and Paul S. Rosenbloom completed doctoral dissertations at Carnegie Mellon University under Allen Newell, introducing SOAR, a cognitive architecture designed to support a broad range of intelligent tasks through a unified problem-space model and a chunking-based learning mechanism.

Blackboard Model Description by Lee Erman, Richard Hayes-Roth, Victor Lesser and D. Raj Reddy

In May 1980, Lee Erman, Richard Hayes-Roth, Victor Lesser and D. Raj Reddy published 'The Hearsay-II Speech-Understanding System: Integrating Knowledge to Resolve Uncertainty' in Artificial Intelligence, vol. 14, providing the canonical description of the blackboard model as a structured framework for cooperative problem-solving among independent knowledge sources.

Augmented Transition Networks Introduced by William A. Woods

In 1970, William A. Woods of Bolt Beranek and Newman published 'Transition Network Grammars for Natural Language Analysis' in Communications of the ACM, introducing Augmented Transition Networks (ATNs) as a formalism for parsing natural language by extending finite-state transition networks with recursion and registers, enabling more expressive grammatical coverage.

SHRDLU Natural Language Understanding Program Developed by Terry Winograd at MIT

In 1970, Terry Winograd at the Massachusetts Institute of Technology completed SHRDLU, a natural language understanding program that allowed a user to converse in English about a simulated world of coloured blocks, demonstrating that a computer could parse and respond to complex grammatical instructions within a constrained domain.

Machine Perception of Three-Dimensional Solids, Lawrence Gilman Roberts (MIT Lincoln Laboratory)

In 1963, Lawrence Gilman Roberts, working at MIT Lincoln Laboratory, completed his doctoral thesis demonstrating that a computer could interpret a 2D photograph of polyhedral objects, reconstruct their 3D structure, and re-render them from arbitrary viewpoints with hidden lines removed, establishing foundational methods for machine interpretation of three-dimensional scenes.