Audrey Tang

I think it’s different. In MoE every token hits different so-called experts, but those experts are still co-pre-trained on the same data bundle — specialised layers in one stack — mainly to cut compute or memory.

鍵盤快捷鍵Keyboard shortcuts

j 下一段next speechk 上一段previous speech