I think it’s different. In MoE every token hits different so-called experts, but those experts are still co-pre-trained on the same data bundle — specialised layers in one stack — mainly to cut compute or memory.
j 下一段next speechk 上一段previous speech