Listen to this lesson
Unlock audio and more
Audio streaming, downloadable PDFs and certificates come with Plus and Pro.
What it means
A text-only model reads and writes text. A multimodal model handles several kinds of data in a shared representation, so it can look at a photograph and describe it, read a chart and answer questions about it, watch a video, or listen to speech and respond in speech.
The technical trick is representing different media in a common space, so an image and the sentence describing it land in comparable territory. That shared representation is what allows genuine cross-modal reasoning rather than a pipeline of separate systems handing off to each other.
Most current frontier models are multimodal to some degree, and the frontier is moving toward natively multimodal training rather than bolting vision onto a text model after the fact.
Why it matters
Multimodality removes the transcription step that used to gate a huge class of real work — screenshots, scanned documents, whiteboard photos, diagrams, recorded meetings. A great deal of business information isn't text, and until recently that meant it was out of reach.
In practice
Check which modalities a model supports for input versus output — they're often asymmetric, and a model that can read images frequently cannot generate them. Also check whether images count toward your context and cost, because they do, often substantially.
Where this shows up
Tools and models in our catalog.
Gemini 3.1 ProGoogle DeepMind's Pro-tier model (Feb 2026) and still the newest Pro-tier Gemini you can use, since Gemini 4 Argon is limited to vetted cyber defenders. 94.3% GPQA Diamond and 77.1% ARC-AGI-2 at launch. 1M token context, native multimodal input (text, image, video, audio), Deep Think reasoning mode. Available via Vertex AI Model Garden and Google AI Studio.
GPT-5.6OpenAI's flagship model family from July 9, 2026 until GPT-6 Astra superseded it on September 3, 2026. Three tiers — Sol (flagship), Terra (balanced), Luna (fastest/cheapest) — each with a ~1.05M context window. GPT-6 Sol and Luna replaced the Sol and Luna tiers on September 22, 2026 at half the price; no GPT-6 Terra has been announced, and GPT-5.6 Luna is still the Free and Go default inside ChatGPT Chat.
Claude Opus 5.5Anthropic's leading model (September 22, 2026) and the first of the Claude 5.5 family — Anthropic says it performs at Claude Fable 5.1's level on most work while costing 40 percent less to run than Opus 5, at $4 input and $20 output per million tokens with cache reads 60 percent cheaper at 20 cents. Ships with Fable 5.1-level biology and cybersecurity safeguards.