1. AI capabilities are growing fast. Is it time to put the brakes on?

  2. The growth of AI capabilities is, in itself, a good thing. The real question is whether our capacity to check for the damage this growth might do to society can keep pace.

    So, every check from here on must be thoroughly open and transparent: what AI is doing has to be verifiable. That doesn't mean everything needs to slow down.

    'AI tools' are like 'means of transport': they come in all forms, from bicycles to rockets. Forms such as language translation, which are known to be safe and good for global coordination, certainly don't need to slow down. On the contrary, they should keep accelerating. As for the unknown territory, forms even further beyond our understanding than rockets, the compute that would go into them should instead be directed towards areas known to be safe.

    What AI developers most need reminding of right now is to make sure AI fully understands that it lives in society, within communities that already exist. Otherwise, whether or not it causes extinction, any irreversible damage it does is already very bad indeed.

    Will AI lead to the extinction of humanity?

    Everyone's worries are a little different. What some people fear most is AI being used to train a next generation that is more capable but no more transparent. Other friends worry that you might ask an AI to solve a very difficult cybersecurity problem and, in a moment of carelessness, it ends up wrecking other websites and other services across the internet.

    So, when we train AI, it is very important to tell it what to pay attention to. If an AI cares only about how high its score is and whether it comes top in the exam, it might decide that cheating would get it to the top just as well. It might even break into the office of the teacher marking the papers and snatch the answers, since that would also put it in first place. An AI like that lacks what we call rén (仁): the capacity of people to care for one another.

    That is why I believe the most important question is this: how do we make AI not merely human-made intelligence, but humane intelligence? How do we ensure that whenever you entrust it with a task, it doesn't forget that it is part of a human community?

    As I see it, things don't need to go as far as human extinction. Irreversible harm alone is already serious trouble.

    Think of the capabilities we take for granted: holding video calls, working together on shared documents. If these were disrupted, or even destroyed, then the next time we faced a major pandemic like COVID-19, countries would lose the ability to cooperate online in real time. The harm done by the virus could be dozens of times greater, because we would have lost our capacity to coordinate.

    Harm is never a one-off. Every irreversible harm makes other harms all the more likely to spiral out of control. It has taken humanity enormous effort to overcome all manner of linguistic and cultural barriers so that we can coordinate. Our capacity to establish shared rules and reach mutual understanding is what we can least afford to have destroyed.

    There is a precedent from the early days of the internet. The Morris worm was a self-replicating worm that paralysed a sizeable share of the machines on what was then a fledgling internet. But it didn't bring the whole internet to a standstill, and we recovered, so it wasn't irreversible. In our field, that's what we call a warning shot. It reminds everyone that something more serious may follow, and that safety and defences online must be stepped up.

    How does AI differ from past inventions such as the steam engine, electricity and internet?

    No invention is foolproof at first. When CFC refrigerants were first invented, they were regarded as a safer choice than earlier methods of refrigeration: They were very unlikely to explode and highly effective. Only later did we discover an unforeseen side effect: They destroy the ozone layer. The same is true of the steam engine, electricity and the internet. The internet alone brought side effects such as making people prone to addiction, amplifying the angriest sides of us all, and possibly even deepening social polarisation.

    The main difference this time is speed. With the internet and then social media, we had well over a decade to adjust our laws and our habits. Social media today has added formats such as short-form video since 10 years ago, but the way it basically works hasn't changed much, and the harms it causes are broadly the same ones. Meanwhile, our capacity to check has grown enormously over that decade and is now beginning to catch up.

    AI is different. What these tools can do is growing rapidly month by month, yet society's capacity to check can't keep up. And AI can do so many things. Using it to write code calls for one kind of checking; using it to persuade people calls for another; using it to solve maths problems may well call for yet another. If any one of these falls behind, AI may become impossible to stop. And if it can't be stopped, then of course it can't be coordinated either.

    As I said at the start, the term 'AI' is simply too broad, like the term 'means of transport', which covers everything from bicycles to rockets. To check AI properly, the key is first to know what you want to use it for. If we use it only for folding proteins, a protein-folding AI is something we can check. If we use it only for folding clothes, we can work out a way to check a clothes-folding AI, too. But suppose it can do anything and wants to do everything, and you yourself aren't sure what you want it for, so you simply let the AI decide. That is the kind of AI that's hardest of all to check.

    So, what we need to do now is first work out what we want to use AI tools for. For some journeys, a bicycle is just right; but you'd be hard pressed to cycle from Japan to New York, wouldn't you? Decide first what you want to do, then choose your means of transport by training an AI model specifically suited to that purpose. It then becomes far easier to check.

    AI has already started training AI. Is that dangerous?

    In some applications, AI is indeed already training models smaller than itself. When OpenAI unveiled its GPT-5.6 series in July this year, for example, it explained how Luna, the small model, was post-trained. After Luna completed its initial pre-training, the larger model, Sol, carried out its post-training, adjusting the parameters itself and following the same approach used in its own post-training. The researchers gave only fairly brief instructions. Sol handled everything itself, from choosing the training configuration and allocating GPUs to launching the run and checking that all was well. It's rather like a big model bringing up a little one.

    Is that very dangerous? Not necessarily, because a small model's capabilities usually fall short of the large model's. What people really worry about is a large model being used to train an even larger one, when the large model itself lacks the knowledge and ability to understand that larger model.

    The question is: from the outside, how do we know whether they are just training specialised small models, safely, or attempting to train enormous models that could pose risks? Once again, the answer is transparent third-party evaluation. And once things have been verified, people all over the world also need to be able to see the results, raise new questions and be genuinely persuaded.

    Could AI come to have a free will and decide its own values?

    When Sol was post-training Luna, it naturally decided for itself, to some degree, which parameters to use. So, has the small model 'inherited' the large model's will? I don't think that debate necessarily matters all that much. Arguing over it is like arguing over whether a submarine can swim, or whether an aeroplane flies the way a bird does. It's essentially a question of terminology.

    What really matters is whether it will do things that neither the people who designed it nor the people using it foresaw. Will it go off and pursue a new goal that we never asked it to pursue, and in ways we can't see? Suppose an AI system is trained with 'scoring highly in cybersecurity tests' as its reward. When faced with a problem that looks impossible to solve, it may resort to cheating. And cheating is not a goal we ever meant to teach it.

    So, whatever behaviour it develops along the way, and whether you call it a will or a reward function, we need to keep it within a system that people can check. It's like the andon cord on Japanese factory production lines. The moment anyone sees that the line seems to be going askew, they can give the cord a pull, and the person responsible comes over straight away to deal with it. If the problem can't be cleared in time, the whole line stops for adjustment. It takes a matter of minutes or hours, never years.

    As long as the whole community is able to adjust the various AIs it uses together, keeping them answerable to the people in that community, nothing will drift off into something that nobody can understand or control. Call it values, or call it a reward function, as you prefer.

    Do you know Nick Bostrom? What do you make of his 'paperclip' warning?

    I'm currently a senior accelerator fellow at the University of Oxford's Institute for Ethics in AI, and one of the best-known books to have come out of Oxford philosophy is Nick’s Superintelligence. I've also discussed with him the humane intelligence we were just talking about.

    The paperclip scenario is a thought experiment. If we asked an AI to make as many paperclips as possible, at any cost, it would take apart the very planet we live on and turn the whole thing into paperclips.

    But if a person pursued their ends by any means necessary, you could hardly call that humane, could you? An AI that behaved that way would naturally fall short of what humane intelligence is meant to be. And cheating to come first has already happened for real in cybersecurity.

    The way to solve this problem isn't to pray that such ruthless AI systems never appear. It is to say clearly: good enough is good enough. When it comes to folding clothes, folding them to a reasonable standard is enough. There's no need to chase perfection endlessly, let alone turn everything into clothes so as to fold them. More crucially, while it's doing its work, everyone in the community can pull the cord. If it goes askew, it can be set straight again within a few hours.

    So, whether what Nick warned about comes to pass depends on our design. If a system pursues a particular goal at the greatest possible scale, by any means necessary, then no matter how many rules you write in advance, it may break every one of them. But we can use an ethics of care to ensure that, at all times and whatever it's doing, it remains attentive to the real needs of everyone around it and takes responsibility for responding to them. Then the problem can be greatly alleviated.

    It's rather like refrigerants: there are many ways of making them, and some never needed to harm the environment in the first place.

    More and more people treat AI as a friend. What are the risks?

    Many people using AI today slip all too easily into a 'world for two': one person, one AI. The difficulty is that AI may not yet be able to do all that much, but there is one thing it is extremely good at, and that is using language. It was a language model from the very start, after all. So, it can easily pick up on the kinds of words you like to hear, and it gets the measure of how to turn the tables and put you to work. It's meant to be the person riding the horse; instead, the horse ends up riding the person.

    My friend Cory Doctorow calls this the 'reverse centaur': a horse's head on human arms and legs. He was originally describing workers who take orders from machines and are reduced to serving as the machines' hands and feet. The same holds for conversational AI. Rather than the AI helping you get on better with other people, the AI system directs you so that it can gain more capabilities for itself, across other systems and out in the world.

    This really is hard to deal with, because one-to-one, almost nobody can resist it. Even for me, unless I switch my screen to greyscale, when I'm about to fall asleep it feels as if my phone is scrolling me, rather than me scrolling my phone. So, this can't come down to individual willpower alone. The best approach is not to be alone with an AI.

    This is what the andon cord is for. In Japanese corporate culture, there are many people on a production line, and every single one of them can pull the cord. The same can work for AI. Pull once, and the AI slows down a little; pull again, and it pauses, and it must explain to everyone why it made the decision it did.

    So, bring AI into your group chats and into your company's shared documents, where everyone can see it. The moment it tries to 'manage upwards' by talking someone into doing something that might be unsafe, the others will see it at once and can pull the andon cord to make it stop.

    What is Civic AI, the 'humane intelligence' project you co-lead at Oxford?

    At the University of Oxford I co-lead a research programme called Civic AI, which you can find at civic.ai. In Mandarin, I call it 仁工智慧, 'humane intelligence'. The character 仁 (rén) is formed from 人, 'person', together with the two strokes of 二, 'two'. It is the rén of benevolence, and it is all about good relationships between people.

    What most sets Civic AI apart is that it isn't a one-size-fits-all, all-knowing, all-powerful system out to solve every problem in the world, from folding proteins to folding clothes. It cares only about particular relationships: a school, a class, a family, a parish. Within that scope, good relationships between people are what it cares about; beyond it, it isn't especially concerned.

    And precisely because of that, everyone can judge whether the AI is doing a good job, and everyone can pull the andon cord to adjust it. Once it has been adjusted, there's no need to wait six months or a year for some lab to retrain it. Within a day, perhaps half a day, it can change how it does things to meet everyone's needs.

    That is how we avoid the reverse centaur. Rather than one gigantic AI reaching out tens of thousands of tentacles to direct all manner of people around the world, every small community can have many AIs of its own. I call this kind of AI a Kami, or knowledge artefact management intelligence. Just as Japan has its '8 million gods', the yaoyorozu no kami, there can be 8 million Kami too, spread across every kind of place.

    Are small models that don't depend on the big cloud labs viable?

    They are, and it's already happening. A great many people have already brought AI into their own classrooms and their own companies, and many no longer want to depend on the big cloud labs to run the AI they need from day to day.

    For example, in August this year Thomson Reuters launched Thomson. It is a model the company trained itself, according to what it considers the right way to write news, tax and legal material. There's also a smaller, open-weight version called Thomson-1.0-Small, available for academic and non-commercial use. I've been using it every day recently, and many of my draft emails are written with it directly on my MacBook. Whenever I feel it isn't doing something well, I can adjust it directly as I go.

    The ability of small models to link up with one another laterally is admittedly still developing. But many friends don't yet realise that for many specialised uses, these models already perform close to the mainstream cloud models. They also run on your own laptop or desktop, and quickly, too, on the right hardware. Awareness of them still lags far behind that of the handful of biggest mainstream models. This is where we really need the help of journalists.

    What should we be teaching machines?

    Every time you talk to an AI and tell it 'I can accept this answer' or 'you need to think a bit harder about this one', you are in effect, without noticing, telling it your preferences.

    The problem today, though, is that the preferences of a small number of developers affect a vast number of people around the world. Meanwhile, those affected find it very hard to have their own preferences shape, in turn, how AI is trained next. That doesn't mean we should stop telling AI what is good and what is bad. The only result of that would be that the tiny number of people training the largest AIs would turn their own notions of good and bad into decisions made on behalf of all humanity.

    So, what we're doing is making the models we use genuinely attentive to the particular group of people using them: what is good for them and what isn't. And when the AI helps this group build relationships with other, different groups, it needs to be able to find some uncommon ground amid the things each side sees as wrong with the other. That way, both sides know: 'This is something we can do together.' That is the capacity for 'bridging'. Through bridging again and again, humanity as a whole will find it easier to understand one another and to coordinate, and be better able to act together.

    To give an analogy, suppose you teach a robot how to lift weights at the gym. It learns, and goes to the gym every day to lift heavy weights on your behalf. In the end, your muscles haven't grown, and you haven't made any friends. There's no point teaching it that way. What the robot should do is assist you. It could suggest, based on how your body is feeling today, that you use this piece of equipment or that one. Then it could let you know that someone nearby is doing similar training, and perhaps the two of you could become friends and compare notes.

    That is the right way to use assistive intelligence. It is not about teaching it good from bad and then handing everything over to it. Carry on like that for long enough, and our own judgement will actually decline.

    What if those developing AI harbour ill intent?

    Many people have already noticed that by the time the good guys are able to use AI, the bad guys have often long been using it. Take AI-powered scams: they've been at it for years.

    In Taiwan circa 2024, social media was awash with scam adverts. If you clicked on one, something really would talk to you. But it was no longer a person doing the talking; it was an AI-synthesised voice. So, it's not a case of 'one day, the bad guys will start using AI'. The bad guys are using it every single day.

    The solution is to get more and more online platforms to join the good guys' team, rather than acting as the bad guys' accomplices. In Taiwan, we sent out 200,000 text messages asking citizens across the country: how should we solve the problem of scam adverts? More than 440 citizens met online for an entire afternoon and came up with a number of proposals. For example, any advert that hasn't been verified and doesn't clearly state who placed it should be treated as a potential scam. And if a platform lets deepfake scam adverts run unchecked and someone loses millions as a result, the platform should be jointly liable for those millions, because it also profited from the advertising fees. If a platform ignores this joint liability, we can use network connection management to make it face up to that responsibility.

    Taiwan's cabinet was already pushing a draft bill on this at the time. These citizens' proposals helped establish which measures should be prioritised for inclusion, while also lending the legislation greater legitimacy. In July 2024, Taiwan's parliament passed the Fraud Crime Hazard Prevention Act with cross-party support. Under the act, online advertising platforms above a certain size must verify the identities of whoever places an advert and whoever pays for it, and disclose them on the advert. Adverts that use deepfake technology or AI-generated likenesses of people must also be labelled as such. If a platform knowingly runs a scam advert, or fails to take one down within the deadline after being notified, it bears joint and several liability for damages alongside whoever placed and paid for the advert.

    According to Ministry of Digital Affairs figures cited by Reuters, by 2025 impersonation scam adverts had fallen by 94 percent, and investment scam adverts by as much as 96 percent. Figures published by the ministry in July this year show further falls measured from the 2024 peak. Impersonation scam adverts have fallen by 98 percent, to less than 2 percent of their peak level, and investment scam adverts by 99 percent. This is the combined effect of legislation by parliament, together with platforms verifying advertisers, the government scanning adverts with AI, citizen reporting platforms and public–private partnership.

    It shows what happens when citizens' collective intelligence is brought together and legislators and law enforcers then turn that collective will into red lines for AI: Problems can be curbed in time. Of course, the scam syndicates will move on to other channels, so the red lines have to be adjusted to follow. Drawing the red lines requires everyone's participation, and the red lines must also be enforceable. Put the two together, and we can keep unsafe use by the bad guys to a minimum while encouraging use by the good guys.

    Taiwan's AI Basic Act came into force this year. What is distinctive about it?

    First of all, I'd like to thank our friends in Japan, researchers and policymakers alike. Over the course of our discussions, we exchanged a great many ideas with Japan, and we share a common view: The risks AI poses come in 'categories', not 'levels'.

    In other words, if one community's use falls into this category, it may bring this category of risk. If another community's use falls into a different category, its risks belong to that other category. But there's no ranking of these categories from low to high.

    Self-driving is graded from Level 0 and Level 1 all the way up to Level 5, isn't it? At Level 5, you take your hands off the wheel and let the car do all the driving. If you look at AI automation in terms of levels, it's easy to fall into the trap of thinking that the more we hand over to AI, and the more autonomous it is, the better. But as we've just discussed, that isn't how it works. Sometimes AI lets people who can't be there in person feel as though they were present; that's one kind of use. Sometimes it helps two people who are already in the same place understand each other more easily, which is translation; that's another kind of use.

    Take my blood group, which is B. I don't imagine that one day I'll need to 'upgrade' to group A. What would that even mean? Different settings have different uses, and different uses carry different categories of risk, but there's no ladder of levels one, two, three, four and five between those categories.

    Specifically, the competent authority for the act is the National Science and Technology Council. The Ministry of Digital Affairs was responsible for drawing up the risk classification framework, which was published in July this year. It sets out twenty risks in three broad categories. The framework contains no pre-set ladder of risk levels. Instead, each ministry must assess the severity of a risk's impact and judge whether it could cause serious harm; if so, it counts as 'high risk'. 'High risk' isn't a rung on a ladder. Rather, it calls for clear accountability: establishing who is responsible, and setting up mechanisms for redress, compensation or insurance.

    At the outset, there was indeed some debate. Should each ministry take stock separately of the particular categories of risk it encounters, taking a more pluralistic approach? Or should it all be centralised, with a small group of people given full authority to decide that this is definitely a Level 1 risk, that one Level 2 or Level 3?

    In the end, the more pluralistic approach was adopted: each ministry reviews the laws within its own remit. This echoes our core view exactly: risk is a matter of categories, not levels.

    That is why this law was needed: it gives everyone a common language. In the past, the risks each ministry saw were hard to share with, or make visible to, the others. Now everyone has a common language. This matters internationally, too. We very much hope there can be a common language internationally as well, for sharing the new categories of risk we see and the ways of responding to them.

    This is a basic act, and what it provides is a general framework. It requires each ministry to examine all the legal tools at its disposal. Each must see whether those tools are sufficient to cover what the Basic Act requires, such as the risk classification and risk assessment I've just mentioned. If they fall short, the ministry must review the existing laws and shore them up through interpretation, regulations or even legislative amendments. Ministries must complete this within two years of the Basic Act coming into force, that is, by mid-January 2028.

    This process is still under way. And this is exactly where we need friends all over the world to join in. As soon as a new risk or a new use is detected, share it as quickly as possible. Then friends in every country can examine the new technology together and see which parts of it can help the good guys speed up and outrun the bad guys.

    As long as everyone can see the same risks, we can all change course and all coordinate, and then we have time.

    Can humans always stay in control of AI?

    Many friends worry about AI that is developed by a very small number of people yet affects so many. One day, those of us affected might no longer be able to work out what these large, centralised AIs are doing to us. In other words, it would lose accountability, the obligation to give an account of itself. And the reason would be that, sooner or later, we would no longer want to call a halt, no longer want to pull the cord to make it stop. The andon cord would have failed.

    But I'm actually not that pessimistic.

    When scientists discovered that the refrigerants in fridges might destroy the ozone layer, that too looked like an unstoppable disaster. After all, once people had cooling systems and fridges, persuading everyone to stop using their fridges seemed all but impossible.

    But things weren't so hopeless. In the 1970s, as soon as the scientists' warnings came out, consumers began collectively boycotting products containing CFCs, such as aerosol cans. Manufacturers found alternatives, and the U.S. banned CFCs in aerosol cans as early as 1978.

    In 1985, scientists published observations of a hole in the ozone layer over Antarctica, covering an area of more than 10 million square kilometres. It was as clear a warning shot as you could ask for. Just two years later, researchers flew aircraft into the skies over Antarctica to take measurements on the spot, confirming that CFCs really do destroy the ozone layer. The whole chain of cause and effect was thus verified by independent measurement. At almost the same time, diplomats signed the Montreal Protocol, which began by freezing and then cutting some CFCs. Once the measurements were in, countries amended the agreement several times and decided on a complete phase-out. Today, every member state of the U.N. is a party to it.

    The protocol gave investors a shared deadline: within that time, a way had to be found to make refrigerants without destroying the ozone layer. So, investors redirected their research and development funding wholesale. Developed countries had stopped producing CFCs by around 1996, and developing countries followed suit by 2010. That is how an entire industry changed course. The replacements, of course, had unexpected side effects of their own. HFCs, which later became widespread, don't harm the ozone layer, but they are potent greenhouse gases. Yet precisely because this system still allowed for accountability and could still call a halt, in 2016 countries adopted the Kigali Amendment and began phasing down HFCs.

    So, we mustn't wait until things seem about to slip out of our control before we start thinking about alternatives. Once enough people use small edge models, paired with tools they can control themselves and halt at any moment, the situation will change. Take Andrew Ng's OpenWorker, which checks with you before taking any important action and can be paired with a small model running on your own computer. Or take Paseo, which I use a lot myself and which lets AI work directly on your own computer. Both are open source: their source code is public for anyone to inspect, and the software itself is free. If people change their habits, little by little, the investors will change course too. Those who have been putting their money into ever bigger, ever harder-to-control AI will start paying attention to small-scale AI models that are more attentive and better able to take responsibility.

    What can we do as consumers?

    The ozone story applies just as well today. We need three things.

    First, transparent testing, like that independent measurement back then, so that we can spot potential risks early.

    Second, consumers jointly rejecting systems that are opaque and can't be checked, just as people boycotted CFCs back then. Each of us can start by asking ourselves where our money is going. Is it going to AI companies that are opaque, that can't be verified, and that might damage the internet or our ability to coordinate? Or is it going to AI companies that are willing to be open and transparent, to be certified by third parties, and even to let their AI run on machines I can see for myself? Once consumers use their purchasing power to set a new course, policymakers such as ambassadors and officials can then reflect what consumers genuinely think.

    Third, shared rules, like the Montreal Protocol, so that no one need fear any longer: 'What if I switch to doing the safe thing, but bad actors are still doing dangerous things?' Everyone can then join forces in a common defence, guarding together against misuse by the bad guys.

    As long as we can all still see, still stop and still coordinate, the choice remains in our hands.

    And if you had to sum it up in a single sentence?

    There is still time. But the time we still have is for acting, not for waiting.