PromptCorrectlyPromptCorrectly
StudioAI CoursesLibraryBlogPricingAbout
Log inStart free
PromptCorrectlyGLOSSARY · HOW MODELS WORK
Home/Glossary/Multimodal AI

What is Multimodal AI?

A multimodal model accepts and/or produces more than text — images, audio, video, documents — so you can prompt with a photo or ask for a picture back.

Text-only models read words. Multimodal models also see images (a screenshot, a chart, a whiteboard photo), hear audio, or generate images and video. That changes what a prompt can be: "what's wrong with this UI?" with a screenshot attached, or "describe the photograph you want" to an image generator.

Prompting principles carry over — be specific, give context, name the format — but each medium has its own vocabulary. Image prompts read like photography briefs; document prompts benefit from telling the model which section matters.

How to use it well
  1. For images in, say what to look at and what to ignore.
  2. For images out, describe subject, setting, lighting, lens and style — and what must not change.
Go deeper
AI image promptsPhoto editing prompts
Related terms
PromptLLM (large language model)All terms →

Use it right now

Ask our brain anything on the homepage — it remembers the whole conversation — or write a brief in the Studio and see the prompt it compiles to.

Ask the brainOpen the Studio
promptcorrectly.comFROM VAGUE INTUITION TO STRUCTURED INSIGHTBYOK · ANTHROPIC · OPENAI · GROK
PromptCorrectlyPromptCorrectly

The visual workspace for people who actually use AI. Built in the open, priced for humans, powered by your own keys.

Product
  • Studio
  • AI Courses
  • Library
  • Pricing
Resources
  • How it works
  • Templates
  • Community library
  • Repository
  • All 3,600 prompts
  • How to prompt correctly
  • Prompts for every field
  • AI glossary
Company
  • About
  • Contact
  • FAQ
  • Pricing
© 2026 PromptCorrectly · From vague intuition to structured insight.
TermsPrivacyRefundDisclaimerAcceptable useCookies