What is Multimodal AI?
A multimodal model accepts and/or produces more than text — images, audio, video, documents — so you can prompt with a photo or ask for a picture back.
Text-only models read words. Multimodal models also see images (a screenshot, a chart, a whiteboard photo), hear audio, or generate images and video. That changes what a prompt can be: "what's wrong with this UI?" with a screenshot attached, or "describe the photograph you want" to an image generator.
Prompting principles carry over — be specific, give context, name the format — but each medium has its own vocabulary. Image prompts read like photography briefs; document prompts benefit from telling the model which section matters.
- For images in, say what to look at and what to ignore.
- For images out, describe subject, setting, lighting, lens and style — and what must not change.
Use it right now
Ask our brain anything on the homepage — it remembers the whole conversation — or write a brief in the Studio and see the prompt it compiles to.