Multimodal models
2 milestones used this technique.
Flamingo: few-shot visual language model for interleaved images, video and text
Researchers introduced Flamingo, a family of Visual Language Models that could handle interleaved images, video and text, achieving state-of-the-art few-shot performance on many benchmarks without task-specific fine-tuning.
SoftBank Robotics and Aldebaran Unveil Pepper, a Humanoid Robot with Emotion Recognition
In June 2014, SoftBank Robotics and its subsidiary Aldebaran Robotics unveiled Pepper, a 1.2-metre humanoid robot equipped with an emotion-recognition system capable of detecting human facial expressions, voice tone, and body language, intended for retail and customer-service deployment.