Updated Sep 10, 2026

Mixture-of-Experts

MoE

An architecture that splits a model into specialized sub-networks and activates only a few per token, cutting the cost of running a very large model.

Share

Listen to this lesson

Free preview · first 0:30
0:00 / 0:30

Unlock audio and more

Audio streaming, downloadable PDFs and certificates come with Plus and Pro.

What it means

In a conventional dense model every parameter participates in every token. A mixture-of-experts model instead contains many sub-networks — experts — plus a small router that picks a handful for each token. Total parameter count can be enormous while the parameters actually used per token stays modest.

This decouples two numbers that used to move together: how much a model knows and how much it costs to run. That is why vendors now quote both total and active parameters, and why a headline parameter count means less than it used to.

The costs are real. All experts must be held in memory even though few are used, so memory requirements track total size while compute tracks active size. Training is also harder — routing can collapse onto a few favored experts unless deliberately balanced.

Why it matters

Most frontier models are now mixtures of experts, which is why capability has kept climbing without inference cost climbing proportionally. It also means comparing models by total parameter count across architectures is meaningless — a large sparse model and a smaller dense one can cost the same to run.

What people get wrong

That the experts are domain specialists. The intuitive picture — one expert for code, one for French, one for medicine — is not what happens. Routing is learned to minimize loss, not to be interpretable, and the specializations that emerge are typically subtle statistical patterns with no correspondence to human categories. You cannot open a mixture-of-experts model and find the legal expert.

That active parameters tell you the hardware you need. They tell you the compute per token; they say nothing about memory. Every expert must be resident even though only a few fire for any given token, so memory scales with the total parameter count, while compute scales with the active count. Sizing a machine on the active number is the most common and most expensive mistake when self-hosting one of these models.

That total parameter count is comparable across architectures. It is not. A large sparse model and a considerably smaller dense one can cost about the same to run, so holding a mixture-of-experts headline number against a dense model's is comparing two different quantities. Read both numbers, or neither.

In practice

When evaluating an open-weight mixture-of-experts model for self-hosting, size your hardware on total parameters, not active. The memory bill is set by what must be resident, not by what fires.

Where this shows up

Tools and models in our catalog.

Related terms