Multimodal
An AI model designed to simultaneously process, understand, and generate data across multiple distinct modalities, such as text, images, and audio.
Think of It Like This
Like a skilled human who can watch a video, listen to a speaker, and read a chart all at the same time to understand a presentation.
Multimodal models project different data types into a shared semantic latent space. This allows a user to upload a photo of a broken machine and verbally ask the model how to fix it. This paradigm is rapidly replacing siloed models, unlocking far more complex and natural human-computer interactions.