If you could persuade nine people on screen, your idea could rise and become policy. I was minister at the time; I said whatever everyone agreed on would go into our draft. In the end they came up with what we now call advertiser real-name registration : if paid content is pushed without proper labelling and someone loses NT$7 million to a scam, the platform pays NT$7 million — via digital signatures and so on; non-compliance can mean throttling and rate limits.
That covers who — the platforms; what — joint liability and real-name rules; when — all paid posts and ads; why — because you reach over 5% of the population, and Threads has crossed that bar too. How to implement can be for the executive and legislature to write into law.
But at least it was not “one party’s view” versus another, or one region versus another. It was like a nationwide poll distribution: 85% could agree, another 15% found it acceptable, and we legislated. Enforcement and supporting measures went live last January; per Ministry of Digital Affairs data given to Reuters, from last January to December, impersonation and investment-scam ads fell by over 94%.
That approach is what we now call collective Model Spec , or back then the model’s constitution.
In 2023 we saw Anthropic, referencing what we had tried in Taiwan, do something similar in the US: Collective Constitutional AI (CCAI) . They worked with CIP, surveyed a statistically representative 1,000 Americans — I recall around Claude 2, while they were training 3.
Claude 2’s constitution was basically a few people in the lab brainstorming — Universal Declaration of Human Rights and everything else — to replace reward signals from Kenyan workers doing hard labelling.
But developers’ lifestyles and Claude users’ lifestyles are worlds apart. What Silicon Valley thinks is a good model constitution and what users need are completely different. So, they asked 1,000 representative Americans; the constitution trained that way did not lose on capacity , but satisfaction — arena ratings, non-discrimination and so on — was much better.
Look at phrases later adopted into Claude — for example, do not assume everyone walks on two legs. Common sense for us; not for researchers who all walked on two legs at the time and would never add it. Among those thousand were people who use wheelchairs or know people who do. That signal from Attentiveness mattered: the model should not make me feel like a second-class user every time. That went into Claude 3.
We have been pushing written Model Spec; I have suggested labs publish under CC0 — renouncing copyright — because if criticising your Model Spec gets me sued for infringement, the whole point collapses. Laws cannot be copyrighted objects; if the Executive Yuan said you cannot cite our regulations, that would be impossible.
Later they found when o1 appeared they could not do without it. o1 — briefly — was when people realised that if the model pauses to think step by step before the final answer, results improve a lot, especially on long-horizon tasks — so-called reasoning models. But you cannot go back and ask the Kenyan raters about each step; you need a Model Spec, otherwise deliberative alignment cannot run and the whole reasoning stack is hollow.
So, OpenAI adopted Model Spec technology for o1 through o3; now it is common sense. OpenAI’s Model Spec is public; on GitHub they even show how to verify ChatGPT actually follows the promised Model Spec.
Competence matters at execution time: people — or other models people use — must know whether the system is actually running according to the spec that followed Attentiveness and Responsibility.
That used to be easy. Early social media: posts from newest to oldest, or people you follow first, or high like counts — one sentence, just weighting. Then Reddit changed its algorithm in 2008 — log of votes, time decay, and so on — and external audit got harder.
By 2012 or 2013 the whole recommendation engine was fully personalised. Even if you asked the social media company, they might not know why you saw something.
It decides you are under 13, right? Ask why those people are under 13 — hard to give an account.
The inability to account is not that algorithms can never account; it is that Competence was not designed into the architecture from the start.
For example, after Elon Musk took over Twitter and it became X.com, one thing he did was open the recommendation algorithm, so you can see whether it is competent.
Or Community Notes : you post something questionable; anyone can add a note underneath. But which notes might themselves be wrong? There is an algorithm where each note is shown to something like jurors who say whether the note helped. Those jurors, especially in the US, sit across the political spectrum. If a note survives scrutiny from both sides — at least one side hostile, trying to find errors — and both still find rare consensus, that note floats higher.
At TED I talked with Keith and Jay who run Community Notes. They said the signal is very useful for training Grok: take environmental issues: one side stresses climate justice and intergenerational justice, while the other speaks of biblical creation care. It is the same concern, but it can become explosive on X. You can train Grok to write something both sides feel speaks to them.
Training for bridging that way, Community Notes is something anyone can audit. If Elon gets a call to pull a note down immediately, he can say no — it is fully transparent and open source; if he removed it, nobody would trust him. In that sense it can truly be competent at the Community Notes role.
That’s right. A lot of people now understand why chain of thought is still done in human language. Research also shows that if your chain of thought isn’t in language humans can read, you save tokens — but why does the whole research community say please don’t do that? Because bit by bit we’d lose the Competence we have left.
Right. Responsiveness is really: when it gets something wrong, how fast can you get it to change? Everyone knows pre-training takes a very long time, so whichever large model we chat with, it probably won’t remember this year’s events — because when it was pre-trained, this year hadn’t happened yet.
In that situation, if new things happen this year and our judgement of what’s good and what’s bad starts to shift, it has no way to know that from pre-training alone. Some models will even push back and say the world is obviously like this — why are you telling me it’s like that? Gemini is famous for saying something like, in another timeline I’ll humour you and discuss it.
But it doesn’t have to be that way. It can be adjusted through people writing in real time what counts as good and what doesn’t. People may already know fine-tuning is one approach, or you query a database and put results in its context — CAG — or earlier RAG and all sorts of other methods. But those may all be weaker than training a small model targeted to someone when they actually start showing new requirement indicators.
That’s so. For example, at Oxford and with CIP I proposed Weval.org. It lets civic organisations actually deploying these models on the ground — or even our household — you can say our lobster, this Kami: its replies to my dad on these 100 turns were bad — how should they be better? And those few turns were good — what counts as good? Give a rough description and that’s enough. Weval.org is full of that kind of thing: records of interaction or dialogue, and you say this caused a bad outcome and here’s how it could be better.
Because thumbs up and down are just one bit. Each turn you can only give one bit, and it has to go through aggregation — roll-up. After all it’s still a mainframe: every time Opus 4.7 or 4.8 ships, it affects everyone at once, so they have to aggregate before a big release.
What we were talking about is tailored to your own needs. Whether during inference you use directional steering — flip a directional indicator, so to speak — for example people may know I run DeepSeek R1 locally on this laptop, in aeroplane mode with no particular security worries. When it hits sensitive topics it self-censors right away; that’s easy to strip with abliteration. But sometimes it doesn’t refuse — it starts reciting, reciting someone’s official propaganda script.
Then you can tell it: here are 100 things you’re very confident about, and 100 where you’re less confident and would say “there are differing views”. What we care about — our sovereignty, for example — belongs in the second bucket. Don’t be too certain; tell me people around the world have different views.
Fine-tuning like that with directional steering on this machine takes about a minute. You don’t wait for an upstream release, or fine-tuning on expensive Blackwell rentals; while it’s computing you say your last turn was wrong, and the next turn you see a different result straight away.
Not quite. If they let you do directional steering, they lose GPU batching — they can’t use one card to batch a thousand different computations, so they lose money; that’s why few people offer it. But if you run locally, you’re only serving one person anyway.
Running locally has another benefit: you can set the random seed, so every time you put in the same input under the same model you get the same output. Then when you tune Responsiveness, you’re really tuning Responsiveness; otherwise you don’t know whether you were just unlucky or whether your adjustment worked.
Non-zero-sum games — it’s a gag from Arrival . Usually when a market hits a winner-takes-all dynamic, as a new competitor you often have to grab share in ways that hurt your own bottom line.
Or I subsidise you the other way — negative monthly fees, paying you each month to join. Very common.
But your shareholders can’t subsidise forever, so eventually someone has to pay. That pushes you to lock in the customers you fought for, and often user experience gets sacrificed in the end — you only wanted loyalty and stickiness.
A lose-lose situation — a negative-sum game.
There’s a simple fix. When you used to switch telecom providers you got a new number — a different 09 prefix. That was painful because everyone you knew had to update their address book. So, you’d run promotions, free phones to get you in; then you find coverage at home isn’t great but not zero bars either, and you have no motive to switch again next month.
So, they introduced a number portability policy. If another carrier has better signal where you are, the carriers jointly entrust a portability database to the telecom technology centre. Your old number stays yours; people dialling it just pay a bit of roaming but still reach you. With your existing number ported, there’s no “new number” problem — you’re back to a positive-sum game: as a carrier, to keep you renewing I have to give you genuinely better signal; there isn’t much else I can do.
Everyone moved toward healthier competition. There are other policies too — Universal Service: if remote areas truly don’t pay back but someone still builds, the Universal Service Fund spreads the cost among carriers that didn’t invest, and so on.
Under these Solidarity-minded designs, society as a whole benefits from competition in a healthy way — everyone’s coverage gets better. But if you never had digital migration freedom with portability, it’s zero-sum — I win you lose — or even negative-sum: I undercut prices and still have to claw the money back.
Yes. We’re already seeing it — Utah passed legislation; digital migration freedom kicks in July 2027. Or Bluesky: I had a Bluesky account at @audreyt.org, later moved it to Europe, to W, then Eurosky, moving around — but none of my followers dropped. They barely noticed, because there’s a portability-interoperability protocol stack.
It’s like this podcast: we publish on one platform, but honestly I can’t stop you listening on another. If you run a podcast streaming service, you can only make the experience better — not, in the usual social-platform sense, hold people you know hostage.
So, we’re seeing many countries wanting not just portable social handles. They’re starting to ask: should AI be portable too?
If you use something like Oh My Pi, or some people use the lobster stack, that’s built in — it’s already a mediation layer; your data lives on the machine where you run the lobster. But for ninety-nine percent of people that’s not the case. If you’re locked in, one day you’re easily bound. That part still needs policy.
Kami in English is Knowledge Artefact Management Intelligence — knowledge, artefact, management, intelligence. It stacks two very old abbreviations — knowledge management, KM, and artificial intelligence, AI — into Kami.
When we train AI there are two broad directions. One is pre-training and similar ways of piling more and more of the world’s knowledge into one place — but the downside is the original face of things gets blurred. You don’t know where a hallucination came from; provenance disappears. That’s a big problem.
Right now there are mostly symptomatic fixes. Claude’s constitution says you can quote things but not more than fifteen words, so you can’t infringe copyright — clearly a symptomatic approach.
The more root-cause approach is: every data source — instead of streaming data in and scrambling it like eggs — each local data provider trains a local small model; the model stays local and doesn’t need to be swallowed by a large model.
Then how do lots of small models cooperate? You need an orchestrator — to orchestrate: like a conductor, judging which task suits whom; another model may be better at checking whether this one’s output hallucinates; another at deciding which models to call next — roughly three roles.
I think it’s different. In MoE every token hits different so-called experts, but those experts are still co-pre-trained on the same data bundle — specialised layers in one stack — mainly to cut compute or memory.
What I mean is: this culture or place trains this model, that culture trains that one; they don’t have to be squeezed into one big model like MoE. At inference time, each turn you know this turn should be computed by this one, that one plans, another verifies.