Updated Sep 10, 2026

Frontier Model

A model at the leading edge of capability, produced by one of the few labs able to run the largest training runs.

Share

Listen to this lesson

Free preview · first 0:30
0:00 / 0:30

Unlock audio and more

Audio streaming, downloadable PDFs and certificates come with Plus and Pro.

What it means

"Frontier" describes position, not quality ranking. A frontier model is one operating at the edge of what has been demonstrated — trained at a scale only a handful of organizations can afford, and capable of things no previously released model could do. The labs producing them are called frontier labs for the same reason.

The practical definition has become partly regulatory. Several frameworks now define the category by training compute above a threshold, on the theory that capability tracks scale closely enough for compute to serve as a workable proxy. That gives the term a legal meaning it did not originally have, and it is why you now see compute thresholds quoted in policy documents.

The set is small and changes hands. Frontier status is a claim about a moment, not a permanent property of a company.

Why it matters

Most AI regulation that targets models rather than uses targets this category specifically — obligations attach above a compute threshold, so whether a model is "frontier" determines which rules apply to whoever built it. For buyers, it is also a rough guide to cost: frontier models are the expensive tier, and a great deal of practical engineering is about not using one where a smaller model would do.

What people get wrong

That "frontier" means "best". It means at the capability edge, from a lab running the largest training runs. Those usually correlate, but the claim is about scale and recency rather than about being the top scorer on your task. A smaller, cheaper, older model routinely beats a frontier model on a narrow job it was tuned for.

That it is a fixed list of companies. Frontier status describes a moment. Labs enter and leave the set as releases land, and a model that was frontier at launch is ordinary within a year or two. Treating it as a permanent brand attribute is how a vendor keeps charging frontier prices for a model that no longer is one.

That the compute threshold is a measure of danger. It is a proxy chosen because it is measurable and hard to fake, not because compute causes harm. A model just under the line is not safe by virtue of being under it, and one just over is not dangerous by virtue of being over. The threshold is a regulatory instrument, not a finding.

In practice

Use the term precisely, or avoid it. When comparing options, ask what a model does on your task rather than which tier it belongs to — and default to the cheapest model that passes your eval rather than the most advanced one available.

Where this shows up

Tools and models in our catalog.

GPT-5.6

OpenAI's flagship model family from July 9, 2026 until GPT-6 Astra superseded it on September 3, 2026. Three tiers — Sol (flagship), Terra (balanced), Luna (fastest/cheapest) — each with a ~1.05M context window. GPT-6 Sol and Luna replaced the Sol and Luna tiers on September 22, 2026 at half the price; no GPT-6 Terra has been announced, and GPT-5.6 Luna is still the Free and Go default inside ChatGPT Chat.

Claude Opus 5.5

Anthropic's leading model (September 22, 2026) and the first of the Claude 5.5 family — Anthropic says it performs at Claude Fable 5.1's level on most work while costing 40 percent less to run than Opus 5, at $4 input and $20 output per million tokens with cache reads 60 percent cheaper at 20 cents. Ships with Fable 5.1-level biology and cybersecurity safeguards.

Gemini 3.1 Pro

Google DeepMind's Pro-tier model (Feb 2026) and still the newest Pro-tier Gemini you can use, since Gemini 4 Argon is limited to vetted cyber defenders. 94.3% GPQA Diamond and 77.1% ARC-AGI-2 at launch. 1M token context, native multimodal input (text, image, video, audio), Deep Think reasoning mode. Available via Vertex AI Model Garden and Google AI Studio.

Grok 4.7

SpaceXAI's flagship for coding and knowledge work, released September 21, 2026 on a new, larger base model at the same $2 input and $6 output per million tokens as Grok 4.6, with a 500,000-token context. On SpaceXAI's own table it leads on legal and electrical-engineering work but trails Claude Fable 5.1 on coding and Terminal-Bench. In the API, Cursor and Grok Build; the Grok app still runs Grok 4.6.

Related terms