混合专家(Mixture of Experts, MoE)

一句话定义:把 Transformer 里那个巨大的前馈网络(FeedForward, FFN)复制成很多份”专家”,每个 token 只用其中少数几个——这样可以在不增加每 token 算力的前提下,把总参数量做大。


1. 它解决什么问题

稠密模型(dense model)的困境:参数量 = 每 token 的算力。想把模型变大,推理就变贵。

MoE 打破了这个等式:参数量可以涨 10 倍,每 token 的算力基本不变。

Switch Transformer 摘要里的定义:

“In deep learning, models typically reuse the same parameters for all inputs. Mixture of Experts (MoE) defies this and instead selects different parameters for each incoming example. The result is a sparsely-activated model — with outrageous numbers of parameters — but a constant computational cost.”

传统模型对所有输入都用同一套参数;MoE 打破这一点,给每个输入挑不同的参数。结果就是——参数量可以夸张地大,但每步的计算成本基本不变。最后那句 “constant computational cost” 是整个设计的目的。

2. 三个核心机制

(1) 路由(routing)+ top-k 选择

“For a token representation x, the router calculates a score for each expert and keeps the top k choices. … Routing is performed for each token. Two neighboring tokens can therefore use different experts.”

关键点:路由是逐 token 的。相邻的两个字可能走完全不同的专家。

(2) 共享专家(shared expert)

总有几个专家对每个 token 都激活,专门承载通用知识:

“This is an expert that is always active for every token. … The benefit of having a shared expert was first noted in the DeepSpeedMoE paper, where they found that it boosts overall modeling performance compared to no shared experts. This is likely because common or repeated patterns don’t have to be learned by multiple individual experts, which leaves them with more room for learning more specialized patterns.”

注意:共享专家的参数同时计入总参数和激活参数。

(3) 负载均衡(load balancing)

如果路由放任自流,会出现路由崩溃(routing collapse)——所有 token 都涌向少数几个专家,其余专家永远学不到东西。

  • 传统解法:加一个辅助损失(auxiliary loss)惩罚不均衡。Switch Transformer 用 α = 10^-2。

    但辅助损失太大会损害性能:“too large an auxiliary loss will impair the model performance”

  • DeepSeek-V3 的解法:无辅助损失负载均衡——给每个专家加一个偏置项,只影响路由选择、不影响门控值;每步训练结束时,过载的专家减 γ、欠载的加 γ(γ = 0.001)。

3. 细粒度专家:把专家切得更碎

DeepSeekMoE论文 的两条策略:

  1. 细粒度专家分割:把专家切成 mN 个、激活 mK 个,激活组合数爆炸式增长。
  2. 共享专家隔离:留 K_s 个专家恒激活。

2026 年的趋势是专家越来越多、每个越来越小:

模型每层专家数每 token 激活
Mixtral 8x7B(2024)82
DeepSeek-V3(2024-12)256 + 1 共享8 + 1 共享
GLM-5(2026-02)256(社区口径)8

4. MoE 到底比稠密强多少(正反证据)

正方(Krajewski et al., arXiv:2402.07871):

“a compute-optimal MoE model trained with a budget of 10^20 FLOPs will achieve the same quality as a dense Transformer trained with a 20× greater computing budget, with the compute savings rising steadily, exceeding 40× when budget of 10^25 FLOPs is surpassed.”

一个用 10^20 FLOPs 算力训练出来的计算最优 MoE 模型,能达到「用 20 倍算力训练的稠密 Transformer」同等的质量;而且这个算力节省会持续扩大,预算超过 10^25 FLOPs 时能超过 40 倍。

反方(同一篇引用的对立研究):

“the gap in efficiency between MoE and standard Transformers narrows at scale (Artetxe et al., 2022) or even that traditional dense models may outperform MoE as the size of the models increases (Clark et al., 2022).”

还有一条挑战常识的:Ludziejewski et al.(arXiv:2502.05172)发现 “MoE models can be more memory-efficient than dense models, contradicting conventional wisdom”——打破了”MoE 一定显存低效”的刻板印象。

5. MoE 的工程代价

问题表现
微调会破坏路由”fine-tuning shifts token distributions. If the router learns too quickly, it can collapse to a subset of experts”
通信密集”MoEs trade dense compute for sparse compute—but replace it with dense communication”
冷启动慢所有专家必须先加载,“slow startup times and expensive autoscaling”
专家容量限制超过容量(expert capacity)的 token 会被直接丢弃,跳过计算走残差连接

6. “专家”这个词会骗人

“The term expert can also be misleading. Experts do not necessarily separate into clean human categories such as mathematics, code, and grammar. Specialization can overlap and may vary across layers or training stages.”

别把专家想象成”数学专家""语文专家”。它们学到的分工是训练过程自发形成的,未必对应人类的知识分类。

相关页面

  • 激活参数 —— MoE 带来的新参数口径
  • 模型参数 —— 参数量到底指什么
  • DeepSeek-V3 —— 最完整的工程样本