Models
Mixture of Experts, Explained: How Sparse Models Work
A mixture of experts activates a slice of a huge model on each token. Here is how routing works, what it costs in memory, and which parameter count matters.
1 article on Emergent Wire tagged "transformers."
1 article
A mixture of experts activates a slice of a huge model on each token. Here is how routing works, what it costs in memory, and which parameter count matters.