Mixture of Experts, Explained: How Sparse Models Work
A mixture of experts activates a slice of a huge model on each token. Here is how routing works, what it costs in memory, and which parameter count matters.
In this story 6 sections
A mixture of experts (MoE) is a model architecture that splits a network into many specialized subnetworks and routes each token to only a few of them. The model holds a large total parameter count but activates a small fraction per token, which cuts the compute cost of both training and inference.
Every frontier lab now ships at least one mixture of experts model, and the reason is arithmetic rather than elegance. Dense models pay for every parameter on every token. MoE models do not.
That arithmetic compounds fast at scale. Serving a trillion queries a day at even a fraction of a cent per token saved adds up to a real line item on an income statement, which is the actual reason the entire industry converged on this architecture within a couple of years.
This explainer covers what a mixture of experts actually does, why routing is the hard part, what the architecture costs in memory and complexity, and how to read the two parameter counts labs publish. It is written for engineers and technically literate readers who want the mechanism, not the metaphor.
What a Mixture of Experts Model Is
Take the feed-forward layer of a transformer and make 64 copies of it. Each copy is an expert. Add a small router network that looks at each incoming token and picks two experts to process it. That is the whole idea.
The result is a model with, say, 400 billion total parameters where any single token touches perhaps 30 billion of them. Training and inference compute scale with the active count. Model capacity scales closer to the total.
The term predates the transformer era by decades. Its modern form traces to sparsely gated layers described in machine learning literature indexed on arXiv, the open-access preprint repository, where most architecture work in this area is published before peer review.
Expert count varies widely between releases. Some models use eight large experts, others use hundreds of small ones, and the choice changes both the routing difficulty and how evenly work spreads across accelerators. There is no consensus optimum.
Experts do not specialize the way the name suggests. Nobody trains a "biology expert." Routing patterns that emerge are usually syntactic or statistical, and they resist clean interpretation.
Researchers who have tried to reverse-engineer what a given expert actually handles usually come back with something underwhelming, like "tokens that follow a comma" or "numeric sequences." The name promises specialization along human-legible lines. The reality is closer to a learned load-balancing scheme that happens to work.
How Routing Works, and Where It Breaks
The router is a small learned layer that scores every expert for a given token and sends the token to the top one or two. Scores are soft, so the model can learn which experts help on which inputs.
Two failure modes dominate training. The first is collapse: the router learns to send almost everything to a handful of experts, and the rest of the network stops learning. The second is thrash, where routing decisions swing between steps and training destabilizes.
The standard fix is an auxiliary loss that penalizes uneven expert utilization. It works, and it costs some quality, because you are partly optimizing for balance rather than for prediction.
The practical consequence: MoE training runs are less forgiving than dense runs. A misconfigured balance coefficient produces a model that trains without errors and underperforms a dense baseline of the same active size.
Distributed training adds another wrinkle. Experts sit on different accelerators, so every routing decision becomes network traffic. At scale, the interconnect, not the math, sets the throughput ceiling.
This is why the largest MoE training runs care as much about the network fabric between accelerators as they do about the accelerators themselves. A faster chip does nothing if the tokens routed to a distant expert are stuck waiting on a slower link.
What a Mixture of Experts Actually Costs
MoE reduces compute per token. It does not reduce memory, and that distinction decides who can run the model.
Every expert has to be resident somewhere, because the router might select it at any moment. A 400-billion-parameter MoE needs roughly the memory of a 400-billion-parameter dense model, while doing the compute of a 30-billion one.
| Dimension | Dense model | MoE model |
|---|---|---|
| Compute per token | All parameters | Active parameters only |
| Memory to serve | All parameters | All parameters |
| Training stability | Well understood | Routing collapse risk |
| Batch efficiency | Predictable | Depends on routing spread |
| Fine-tuning | Straightforward | Harder, router drifts |
That memory profile is why MoE suits a datacenter and not a laptop. The same tradeoff runs in the opposite direction for small models built to run on device, where total footprint is the binding constraint.
A team choosing between the two architectures for a specific deployment is really choosing between two different bottlenecks. Dense and small favors a device with limited memory. MoE and large favors a datacenter with cheap memory and an interconnect built for exactly this pattern.
How Do You Read the Parameter Counts Labs Publish?
Reading a model announcement means finding two separate numbers, not one. Total parameters describe capacity and how much memory the model needs to serve. Active parameters per token describe compute cost, and roughly predict latency and price per million tokens.
A release that publishes only the total is telling you the flattering number. At Emergent Wire we treat an undisclosed active count as a signal that the comparison being invited is not the useful one.
The gap matters for benchmark comparisons too. A 400B total, 30B active MoE compared against a 30B dense model is not a like-for-like test, and neither is comparing it against a 400B dense model. Both framings are common, and both flatter someone. Evaluation practice here is uneven, which is a broader problem covered in our guide to reading AI benchmark scores.
Compute trends by model, including training costs and parameter counts, are tracked by Epoch AI, a research group that maintains public datasets on machine learning models. Their data is the most consistent public source for comparing scale across labs and years.
Why the Mixture of Experts Design Shapes What You Can Run
For anyone serving models, MoE changes the deployment math in three specific ways.
- Cost per token drops, hardware requirements do not. You still need enough accelerator memory for the full model.
- Latency becomes less predictable. Routing spread varies by input, and tail latency widens compared to a dense model of similar quality.
- Fine-tuning gets harder. Adapting an MoE risks disturbing the router, so many teams stick to prompting or lightweight adapters.
Quantization interacts badly with sparsity in some setups. Compressing weights to four bits saves memory across all experts, but rarely used experts degrade first, and the damage shows up on exactly the inputs that needed them.
That failure mode is easy to miss in testing, since it only shows up on the tail of inputs a rarely used expert actually handles. A general quality benchmark rarely samples that tail deeply enough to catch the degradation before it reaches production.
Serving economics drive the adoption. The compute savings show up directly in the price of a million tokens, which is the number that decides whether an application is viable at scale. That pressure is the same one behind the infrastructure buildout tracked in our reporting on the power constraints under AI datacenters.
Whether MoE is the permanent answer is unsettled. Dense models remain easier to train, easier to fine-tune, and easier to reason about, and several labs still ship dense flagships. Emergent Wire reads the current split as an economics decision rather than a settled architectural verdict.
Expect that split to keep shifting as accelerator memory gets cheaper relative to compute. If memory stops being the constraint that makes MoE attractive, dense architectures could reclaim ground simply because they are the easier system to operate, not because they became better at the underlying task.
The Short Version
A mixture of experts trades memory for compute. You hold every parameter in memory and pay for only a few per token, which makes large models cheaper to serve without making them cheaper to store.
When you read the next model announcement, find the active parameter count first. It tells you more about what the model will cost you than the headline number does. Emergent Wire covers model architecture with that number in mind because it is the one that reaches the invoice.
Emergent Wire covers AI models, capabilities, and the industry building them.