Models

Mixture of Experts, Explained: How Sparse Models Work

A mixture of experts activates a slice of a huge model on each token. Here is how routing works, what it costs in memory, and which parameter count matters.

Elena Vasquez

Former ML Researcher, Industry Analysis Lead

Published 5 min read
Elegant 3D visualization of neural networks showcasing abstract connections in a digital space.
In this story 6 sections

A mixture of experts (MoE) is a model architecture that splits a network into many specialized subnetworks and routes each token to only a few of them. The model holds a large total parameter count but activates a small fraction per token, which cuts the compute cost of both training and inference.

Every frontier lab now ships at least one mixture of experts model, and the reason is arithmetic rather than elegance. Dense models pay for every parameter on every token. MoE models do not.

That arithmetic compounds fast at scale. Serving a trillion queries a day at even a fraction of a cent per token saved adds up to a real line item on an income statement, which is the actual reason the entire industry converged on this architecture within a couple of years.

This explainer covers what a mixture of experts actually does, why routing is the hard part, what the architecture costs in memory and complexity, and how to read the two parameter counts labs publish. It is written for engineers and technically literate readers who want the mechanism, not the metaphor.

What a Mixture of Experts Model Is

Take the feed-forward layer of a transformer and make 64 copies of it. Each copy is an expert. Add a small router network that looks at each incoming token and picks two experts to process it. That is the whole idea.

The result is a model with, say, 400 billion total parameters where any single token touches perhaps 30 billion of them. Training and inference compute scale with the active count. Model capacity scales closer to the total.

The term predates the transformer era by decades. Its modern form traces to sparsely gated layers described in machine learning literature indexed on arXiv, the open-access preprint repository, where most architecture work in this area is published before peer review.

Expert count varies widely between releases. Some models use eight large experts, others use hundreds of small ones, and the choice changes both the routing difficulty and how evenly work spreads across accelerators. There is no consensus optimum.

Experts do not specialize the way the name suggests. Nobody trains a "biology expert." Routing patterns that emerge are usually syntactic or statistical, and they resist clean interpretation.

Researchers who have tried to reverse-engineer what a given expert actually handles usually come back with something underwhelming, like "tokens that follow a comma" or "numeric sequences." The name promises specialization along human-legible lines. The reality is closer to a learned load-balancing scheme that happens to work.

Detailed image of illuminated server racks showcasing modern technology infrastructure.

How Routing Works, and Where It Breaks

The router is a small learned layer that scores every expert for a given token and sends the token to the top one or two. Scores are soft, so the model can learn which experts help on which inputs.

Two failure modes dominate training. The first is collapse: the router learns to send almost everything to a handful of experts, and the rest of the network stops learning. The second is thrash, where routing decisions swing between steps and training destabilizes.

The standard fix is an auxiliary loss that penalizes uneven expert utilization. It works, and it costs some quality, because you are partly optimizing for balance rather than for prediction.

The practical consequence: MoE training runs are less forgiving than dense runs. A misconfigured balance coefficient produces a model that trains without errors and underperforms a dense baseline of the same active size.

Distributed training adds another wrinkle. Experts sit on different accelerators, so every routing decision becomes network traffic. At scale, the interconnect, not the math, sets the throughput ceiling.

This is why the largest MoE training runs care as much about the network fabric between accelerators as they do about the accelerators themselves. A faster chip does nothing if the tokens routed to a distant expert are stuck waiting on a slower link.

Detailed view of a microchip on a printed circuit board, showcasing electronic components.

What a Mixture of Experts Actually Costs

MoE reduces compute per token. It does not reduce memory, and that distinction decides who can run the model.

Every expert has to be resident somewhere, because the router might select it at any moment. A 400-billion-parameter MoE needs roughly the memory of a 400-billion-parameter dense model, while doing the compute of a 30-billion one.

What a Mixture of Experts Actually Costs
DimensionDense modelMoE model
Compute per tokenAll parametersActive parameters only
Memory to serveAll parametersAll parameters
Training stabilityWell understoodRouting collapse risk
Batch efficiencyPredictableDepends on routing spread
Fine-tuningStraightforwardHarder, router drifts

That memory profile is why MoE suits a datacenter and not a laptop. The same tradeoff runs in the opposite direction for small models built to run on device, where total footprint is the binding constraint.

A team choosing between the two architectures for a specific deployment is really choosing between two different bottlenecks. Dense and small favors a device with limited memory. MoE and large favors a datacenter with cheap memory and an interconnect built for exactly this pattern.

Colorful data visualization of stock market trends with financial charts.

How Do You Read the Parameter Counts Labs Publish?

Reading a model announcement means finding two separate numbers, not one. Total parameters describe capacity and how much memory the model needs to serve. Active parameters per token describe compute cost, and roughly predict latency and price per million tokens.

A release that publishes only the total is telling you the flattering number. At Emergent Wire we treat an undisclosed active count as a signal that the comparison being invited is not the useful one.

The gap matters for benchmark comparisons too. A 400B total, 30B active MoE compared against a 30B dense model is not a like-for-like test, and neither is comparing it against a 400B dense model. Both framings are common, and both flatter someone. Evaluation practice here is uneven, which is a broader problem covered in our guide to reading AI benchmark scores.

Compute trends by model, including training costs and parameter counts, are tracked by Epoch AI, a research group that maintains public datasets on machine learning models. Their data is the most consistent public source for comparing scale across labs and years.

A software developer working on code at a dual monitor setup in a modern office.

Why the Mixture of Experts Design Shapes What You Can Run

For anyone serving models, MoE changes the deployment math in three specific ways.

  1. Cost per token drops, hardware requirements do not. You still need enough accelerator memory for the full model.
  2. Latency becomes less predictable. Routing spread varies by input, and tail latency widens compared to a dense model of similar quality.
  3. Fine-tuning gets harder. Adapting an MoE risks disturbing the router, so many teams stick to prompting or lightweight adapters.

Quantization interacts badly with sparsity in some setups. Compressing weights to four bits saves memory across all experts, but rarely used experts degrade first, and the damage shows up on exactly the inputs that needed them.

That failure mode is easy to miss in testing, since it only shows up on the tail of inputs a rarely used expert actually handles. A general quality benchmark rarely samples that tail deeply enough to catch the degradation before it reaches production.

Serving economics drive the adoption. The compute savings show up directly in the price of a million tokens, which is the number that decides whether an application is viable at scale. That pressure is the same one behind the infrastructure buildout tracked in our reporting on the power constraints under AI datacenters.

Whether MoE is the permanent answer is unsettled. Dense models remain easier to train, easier to fine-tune, and easier to reason about, and several labs still ship dense flagships. Emergent Wire reads the current split as an economics decision rather than a settled architectural verdict.

Expect that split to keep shifting as accelerator memory gets cheaper relative to compute. If memory stops being the constraint that makes MoE attractive, dense architectures could reclaim ground simply because they are the easier system to operate, not because they became better at the underlying task.

The Short Version

A mixture of experts trades memory for compute. You hold every parameter in memory and pay for only a few per token, which makes large models cheaper to serve without making them cheaper to store.

When you read the next model announcement, find the active parameter count first. It tells you more about what the model will cost you than the headline number does. Emergent Wire covers model architecture with that number in mind because it is the one that reaches the invoice.

Emergent Wire covers AI models, capabilities, and the industry building them.

What does mixture of experts mean in AI models?
A mixture of experts model splits its feed-forward layers into many parallel subnetworks called experts, then uses a small router to send each token to only one or two of them. The model has a large total parameter count but activates a small share per token, cutting compute cost.
Does a mixture of experts model use less memory?
No. Every expert must stay loaded because the router can select any of them at any moment, so serving memory tracks total parameters. What drops is compute per token, which lowers cost and latency. MoE trades memory for compute rather than reducing both.
What is the difference between total and active parameters?
Total parameters describe the model's full size and determine how much memory serving requires. Active parameters describe how many are used on a single token, which drives compute cost, latency, and price. A model can have 400 billion total and 30 billion active.
Why do mixture of experts models fail during training?
The most common failure is router collapse, where the router learns to send nearly all tokens to a few experts and the rest stop learning. Training adds an auxiliary load-balancing loss to prevent this, which costs some quality in exchange for stability.
Do experts in an MoE model specialize in specific topics?
Rarely in a way humans would recognize. Despite the name, routing patterns that emerge during training tend to be syntactic or statistical rather than topical. You do not get a biology expert and a math expert, and interpreting what any single expert handles is difficult.