Listen to this lesson
Audio & video lessons are paid features
Plus unlocks audio streaming. Pro adds downloadable audio, video, certificates, and more.
- Audio streaming
- Downloadable PDFs
- All AI Playbooks
- Personalized content
- Certificates of completion
- Audio MP3 downloads
- Video lessonssoon
- & More…soon
Watch this lesson
What it means
"Frontier" describes position, not quality ranking. A frontier model is one operating at the edge of what has been demonstrated — trained at a scale only a handful of organizations can afford, and capable of things no previously released model could do. The labs producing them are called frontier labs for the same reason.
The practical definition has become partly regulatory. Several frameworks now define the category by training compute above a threshold, on the theory that capability tracks scale closely enough for compute to serve as a workable proxy. That gives the term a legal meaning it did not originally have, and it is why you now see compute thresholds quoted in policy documents.
The set is small and changes hands. Frontier status is a claim about a moment, not a permanent property of a company.
Why it matters
Most AI regulation that targets models rather than uses targets this category specifically — obligations attach above a compute threshold, so whether a model is "frontier" determines which rules apply to whoever built it. For buyers, it is also a rough guide to cost: frontier models are the expensive tier, and a great deal of practical engineering is about not using one where a smaller model would do.
What people get wrong
That "frontier" means "best". It means at the capability edge, from a lab running the largest training runs. Those usually correlate, but the claim is about scale and recency rather than about being the top scorer on your task. A smaller, cheaper, older model routinely beats a frontier model on a narrow job it was tuned for.
That it is a fixed list of companies. Frontier status describes a moment. Labs enter and leave the set as releases land, and a model that was frontier at launch is ordinary within a year or two. Treating it as a permanent brand attribute is how a vendor keeps charging frontier prices for a model that no longer is one.
That the compute threshold is a measure of danger. It is a proxy chosen because it is measurable and hard to fake, not because compute causes harm. A model just under the line is not safe by virtue of being under it, and one just over is not dangerous by virtue of being over. The threshold is a regulatory instrument, not a finding.
In practice
Use the term precisely, or avoid it. When comparing options, ask what a model does on your task rather than which tier it belongs to — and default to the cheapest model that passes your eval rather than the most advanced one available.
Where this shows up
Tools and models in our catalog.
GPT-5.6OpenAI's flagship model family, generally available July 9, 2026 across ChatGPT, Codex, and the API. Three durable tiers — Sol (flagship), Terra (balanced), Luna (fastest/cheapest) — each with a ~1.05M context window, plus an ultra mode that coordinates subagents. Reports state-of-the-art Terminal-Bench 2.1 (88.8%, ultra 91.9%) and Agents' Last Exam (53.6), but all scores are vendor-reported and no SWE-bench Pro number is published (where Claude Fable 5 led the prior generation).
Claude Opus 5Anthropic's near-frontier flagship (July 2026) — close to Claude Fable 5's intelligence at half the price and number one on Artificial Analysis at launch, with token efficiency as its headline: comparable results in fewer tokens and fewer turns than Opus 4.8.
Gemini 3.1 ProGoogle DeepMind flagship model (Feb 2026). 94.3% GPQA Diamond (highest ever), 77.1% ARC-AGI-2, #1 on 12+ benchmarks. 1M token context, native multimodal input (text, image, video, audio), Deep Think reasoning mode. Available via Vertex AI Model Garden and Google AI Studio.
Grok 4.6xAI's flagship model, released August 12, 2026 and built for long-running agents. Scores 61 on the Artificial Analysis Intelligence Index — level with GPT-5.6 Sol Max — at $2 input and $6 output per million tokens, leading the frontier on legal and professional-domain benchmarks while trailing on Terminal-Bench.