
Later they found when o1 appeared they could not do without it. o1 — briefly — was when people realised that if the model pauses to think step by step before the final answer, results improve a lot, especially on long-horizon tasks — so-called reasoning models. But you cannot go back and ask the Kenyan raters about each step; you need a Model Spec, otherwise deliberative alignment cannot run and the whole reasoning stack is hollow.