Listen to this lesson
Unlock audio and more
Audio streaming, downloadable PDFs and certificates come with Plus and Pro.
What it means
Reasoning models are trained to generate an extended internal working-out before producing a final answer, and to do so without being asked. On problems with verifiable answers — mathematics, competitive programming, logical puzzles — this produces large gains over models that answer immediately.
The shift this represents is worth naming. Historically, better performance meant a bigger model trained on more data. Reasoning models improve by spending more compute at inference time instead, which is a genuinely different scaling axis and one that costs the operator on every request rather than once at training.
The tradeoff is direct: slower and more expensive per query. Many vendors now expose the effort level as a control, and the honest framing is that reasoning models are for hard problems, not all problems.
Why it matters
This is the axis frontier competition currently runs along, and it changes the cost model — a reasoning model can consume many times the tokens of a standard one for the same visible answer, which is invisible on a bill until you look at token counts.
In practice
Route by difficulty rather than defaulting everything to a reasoning model. Summarizing an email doesn't need it; debugging a subtle logic error does.
Where this shows up
Tools and models in our catalog.
DeepSeek R1First open-source reasoning model matching OpenAI o1. MIT license. R1-0528 adds JSON output and function-calling. Distilled variants 1.5B-70B. Banned on gov devices in multiple countries.
GPT-5.6OpenAI's flagship model family from July 9, 2026 until GPT-6 Astra superseded it on September 3, 2026. Three tiers — Sol (flagship), Terra (balanced), Luna (fastest/cheapest) — each with a ~1.05M context window. GPT-6 Sol and Luna replaced the Sol and Luna tiers on September 22, 2026 at half the price; no GPT-6 Terra has been announced, and GPT-5.6 Luna is still the Free and Go default inside ChatGPT Chat.
Claude Opus 5.5Anthropic's leading model (September 22, 2026) and the first of the Claude 5.5 family — Anthropic says it performs at Claude Fable 5.1's level on most work while costing 40 percent less to run than Opus 5, at $4 input and $20 output per million tokens with cache reads 60 percent cheaper at 20 cents. Ships with Fable 5.1-level biology and cybersecurity safeguards.
Gemini 3.1 ProGoogle DeepMind's Pro-tier model (Feb 2026) and still the newest Pro-tier Gemini you can use, since Gemini 4 Argon is limited to vetted cyber defenders. 94.3% GPQA Diamond and 77.1% ARC-AGI-2 at launch. 1M token context, native multimodal input (text, image, video, audio), Deep Think reasoning mode. Available via Vertex AI Model Garden and Google AI Studio.