AI Glossary
What does multimodal mean in AI?
Multimodal: A multimodal model understands more than text — images, audio, sometimes video — in one system, and can often produce them too.
In plain English
Early LLMs were brilliant pen pals: text in, text out. Multimodal models grew eyes and ears. Show one a photo of your error screen, a hand-drawn sketch, or a restaurant menu in another language, and it works with what it sees — no typing required.
Why it matters
- Photograph a form and have it filled; screenshot a chart and have it explained
- Voice conversations with AI are multimodality in action
- Collapses many separate tools (OCR, image captioning, speech-to-text) into one
A concrete example
Snap a photo of your electricity bill and ask "why is this 40% higher than last month?" — the model reads the line items straight from the image.
Related terms
Learning AI from scratch? Start with the free Learn AI course picks, then put it to work with 50 tested prompts — or get the daily 2-minute briefing, free →