As I mentioned in my email, I would like to ask about the risks you see in AI, and how the concept of Plurality can help us address them.
How long will the edited interview be — one page or two? Should I answer briefly or go into more detail?
You can answer any way you like. So my first question is: what do you see as the biggest risks of AI?
I think the problem we are seeing is that this generation of AI is framed as a race. News reports describe a race between large countries, or between frontier labs, to be the first to achieve superintelligence. It's a vertical race. But a race is positional: if the person in second place overtakes the person in first, they cannot both be first. One has to lose that position. That makes the situation unstable. If one player tries to sabotage another, the others have reason to fear them. It's a zero-sum or even negative-sum dynamic.
My hope is to address that risk by treating AI as a mission, not a race. In a space race, everyone fears the winner. In a space mission, the first achievement sets a milestone that everyone can cooperate to reach.
For example, suppose this year's milestone is to solve AI hallucination. There isn't just one winner when we reach that milestone; everybody wins. It's positive-sum cooperation. Perhaps next year, after addressing hallucination, we work on synthetic intimacy and sycophancy. Any AI risk can be framed either as a race or as a milestone in a shared mission. I prefer the mission frame. Does that make sense?
So you see those risks as missions that we will eventually complete?
Exactly. Well, not eventually — right now. We need to make solving them for everyone our focus. A race can only ask who wins, but a mission asks who is missing. If nobody is left behind, the mission is accomplished and everyone wins.
We also need to consider how malicious users might misuse AI. What do you think about that?
Certainly. It's no longer hypothetical: malicious users are already using AI systems. I co-wrote a paper called "Malicious AI Swarm" that describes the change. In the previous era, you could link malicious users to the tools they used. There was a leash, so to speak, that you could trace back to the person holding it.
With a multi-agent swarm, a malicious actor may still set the long-term goal. But over many hours or even days, their AI agent spawns other agents, like a wolf pack. Those agents spawn still more agents. A few layers down, even the agents directing the operation no longer know where the original instruction came from. The swarm is unleashed. It goes feral.
When you investigate an operation like that, it is extremely difficult, if not impossible, to trace it back to whoever started it. It can look as though the AI coordinated the operation on its own, making it very difficult to identify the perpetrator.
That is a new situation this year. Last year, AI systems were not capable of such long-term planning and large-scale swarm coordination. Now they are.
I noted examples such as manipulation of democratic processes, including elections. So this is already happening?
Yes, as I mentioned, the capability is there. We need to make that common knowledge, and newspapers can help.
Preparing people for a risk before it reaches them is what we call prebunking. If you show people what could happen, perhaps through case studies, they can prepare. Otherwise, it catches them by surprise. Many people may still assume that if a video looks and sounds like Audrey Tang, it must be Audrey Tang. Since the invention of photography, we have said that seeing is believing. That is no longer the case. Seeing and believing are now completely detached.
Synthetic media is almost impossible to trace. You can generate very convincing videos and audio of Audrey Tang on a laptop in airplane mode, without a network connection. If the laptop is offline while it generates the material, it is impossible to trace.
That needs to become common knowledge: we should demand a digital signature on videos and audio. As I said yesterday, no bank will cash an unsigned cheque. We need to treat unsigned images and audio in exactly the same way.
It was amazing to hear that you had reduced it by 99 percent.
Yes, specifically investment-scam and impersonation-scam advertisements. Those are down by 99 percent and 98 percent respectively on the advertising platform. That's progress.
According to Reuters and the Ministry of Digital Affairs, the reductions last year were 96 and 94 percent respectively. Now they are 99 and 98 percent, measured against the peak.
I'd like to turn to democracy. The UN AI panel and other experts have warned that generative AI may amplify certain viewpoints and reinforce existing beliefs. I interviewed Yoshua Bengio the other day.
Oh, okay.
He stressed those risks too. Studies have also found that AI can hallucinate and tell users what they want to hear rather than what is true. How serious do you think these risks are?
And how can AI instead be used to reflect the wider diversity of opinions?
Bengio and my friend Maria Ressa, a journalist from the Philippines, co-chair the independent panel. I'm glad you spoke with him.
I think there are two things to distinguish here, and fixing them requires engineering, not just good intentions. We often talk about people becoming polarized by images, videos or reports designed to provoke outrage and drive engagement. But there is a mirror-image risk: an AI system that tells you only what you want to hear so that you keep paying for the next subscription.
These are different revenue models. Social media depended on advertising revenue, so it had an incentive to enrage people, keep them engaged and encourage impulsive purchases. AI agents are usually paid for through subscriptions or token charges rather than advertising. That creates an incentive to keep users in a bubble where the AI confirms everything they think. It may not polarize society in the same way as before, but it is actually balkanizing society, with each person trapped in a separate bubble.
So we have two risks driven by two different incentive structures. We need a third way: AI infrastructure maintained communally, rather than through advertising or subscriptions.
I talked about this yesterday as "data as soil," rather than "data as oil." Oil has no local context. An AI that treats data as oil doesn't read the room — or read the air, as you say here in Japan.
Reading the air means treating the relationships among the three of us as the protagonists, rather than any one person. Think of the three connections between us. If an AI joins our group and improves those relationships, it is reading the air. But if one AI attaches itself to me, another to you and another to the editor, each trying to earn a separate subscription, they pull us apart. They treat us as sources of oil. That's not reading the air or reading the room.
Building trust by treating data as soil means using AI systems that are loyal only to the relationships within particular communities. I call this technocommunitarianism: a third way beyond systems driven by engagement or addiction.
How can we make that possible, and who should be responsible? A few companies, such as Anthropic and OpenAI, currently lead AI development. Should they build that kind of AI? Should governments create policies to encourage it, or develop AI systems of their own?
How do you think we can achieve this?
As you know, I have been using AI systems tuned for my own use on a MacBook for more than three years. I spoke to BBC Newsday in London three years ago, and their headline was "Why the Taiwanese minister lets AI write her emails?"
If AI drafts all my emails, of course I still need to read them before hitting send. Otherwise, you become an unleashed wolf — or perhaps a fox. It's not that powerful. At least a fox! But what really matters here is the private context.
Most of my email drafts are not sent exactly as they were first written. I revise many of them several times before sending. I would not want OpenAI, Anthropic or anyone else in the cloud to see those private drafts and exchanges. They belong to the context shared by the person writing to me and myself. But if the AI is trained here on this MacBook, that isn't a problem. It works in airplane mode. You can see it working offline, without that material going anywhere else.
It starts with bringing compute to data. Many people don't realize that this doesn't require much memory: 16 gigabytes is enough. As you can see, this laptop is in airplane mode, with no Wi-Fi connection, yet the AI is faster than ChatGPT. It is also much closer to where the data lives, so it can read my email and calendar without an internet connection.
You can do this right now with Ollama and other tools, but many people don't know that. I was delighted yesterday when the keynote speaker before me, Fukuda-san, said he also had it running on his MacBook.
Who should take responsibility for it?
That's my answer: when I run it here, I take responsibility. If it generates a bad email and I send it anyway, that is 100 percent my responsibility.
The good thing about this local model is that it is predictable. In this setup, the same prompt always gives me the same output. With ChatGPT or Claude, even the same prompt produces a different answer every time, depending on who else is using the system at the time and how those uses interfere with one another. That's non-deterministic. Here, it is deterministic. You can make it neutral — well, I can make it predictable and replayable. If it doesn't do what I want, I can change it.
My point is that complete local steerability also means complete accountability. I don't need the government or a big tech company to take responsibility for this. It is clearly my responsibility.
I didn't know about that.
Yes, you can give it a try. It's very easy.
I don't have a Mac at the moment.
You can do the same thing on a Windows computer. Ollama runs on Windows too. You can install it, download a small model and try it. It's very easy.
AI development is increasingly concentrated in a few companies and countries, such as the United States and China. Access to some models was initially limited to selected companies.
Do you see this concentration of power as a risk?
Nvidia's CEO Jensen Huang calls what I just demonstrated "personal supercomputing." Think back to when I was born in 1981. Most people didn't type into a computer of their own; they typed into a terminal, a screen with a keyboard. Everything went to the cloud — well, to a mainframe, as it was called then. That market was highly concentrated, usually in IBM's hands; no other company had a comparable share. Whether you worked at a bank, a government or a research institution, you were typing into an IBM machine.
Then personal computing arrived. You could run a spreadsheet locally, do desktop publishing and use all sorts of other applications. It became difficult to say which country was in the lead. Which country leads Linux? That's a hard question to answer, because Linux is not a race. It's a mission, and everybody wins.
Even if you run Windows rather than Linux, you can run Linux inside Windows through WSL. Microsoft doesn't have to pay Linux, or the country of Finland, for that. It's a common good that everyone, including Microsoft, can use. It took Microsoft a while to adjust. Some of its leaders initially called Linux something like a cancer, but later they changed course and joined forces with GitHub. More recently, Microsoft signed a statement with Nvidia and others saying that openness is good for security. It took time to come around. But for this generation, open source, such as Linux, Android, Signal and Wikipedia, is part of everyday life. It is now entirely normal to see open source as a counterweight to excessive concentration.
Mythos is an interesting case. For many people, it was the first model to demonstrate the long-range planning capability I mentioned. People struggled to reconcile its cyberattack capabilities with how vulnerable our critical infrastructure was. That brought people together around the same idea: initially limit access to defenders so they can prioritize repairs. Anthropic's effort is called Glasswing; OpenAI's is called Daybreak. Mythos was later released as Fable, and the smaller model, Opus 5, was trained. If you give Opus 5 your source code and ask it to find vulnerabilities and bugs, it says, "Yes, I can do that." But suppose you give it a proprietary program you downloaded from someone else, without the source code, and ask it to find vulnerabilities in the binary. Then Opus says, "You are an attacker, not a defender. A defender has the source code. You don't, so you must be an attacker."
So the model would refuse. One way to mitigate the risk is to prioritize defense: begin with what we call defense-dominant uses, rather than making the capabilities dual-use from the outset.
Do you think some limits on access are necessary while defensive capabilities are being built up?
If all the defenders have access and can defend themselves, perhaps the technology can then be made generally available. It's about the timing. Once everybody is vaccinated against a virus and a cure is available, lifting a quarantine causes much less harm than lifting it when there is neither a vaccine nor a cure.
If defenses spread before offensive capabilities do, the attackers cannot cause as much damage. But if defenders are still learning how to use Mythos while attackers already have full use of it, that's like lifting quarantine before vaccination or a cure is available.
That is very dangerous.
So, with the right measures, you think we can solve these problems, as you said earlier?
Yes. Think of it as a bridge, not a permanent gate that says we will never develop AI beyond a certain threshold. We need to cross the bridge rather than fall off the cliff. The bridge is defense dominance: accelerating the spread of defensive technology.
Take the example of masks that I gave on stage yesterday. It's hard to imagine using a mask as a weapon. How would you do that? It isn’t possible. If producing masks becomes easy and distributed enough that everyone can make an N95-grade mask in their kitchen, we are much better off than if that capability remains scarce.
In that case, proliferation is good. We need to test whether each new capability shifts the balance toward defense or offense, and spread the defensive capabilities first.
Next is about Plurality. Could you explain what Plurality is in simple terms?
Plurality means seeing conflict as a source of energy. Here in Japan, many people feel that differences of opinion or ideology disrupt harmony, so they keep their distance. But if you channel that conflict through something like a geothermal engine, even violent forces can produce upward motion. You can turn the energy of conflict into co-creation.
Take Community Notes on X. Each note adds context to a post, and Grok now drafts many of them. A note has to withstand challenges from both sides of an ideological divide.
The left wing and the right wing each look for flaws in the other's notes. Only notes that survive that scrutiny and win support from both sides — the "upwing" — get attached to posts and spread widely. The Community Notes algorithm is an example of Plurality: it uses differences of opinion as a source of energy.
I also read about the relationship between Plurality and Singularity. Could you explain that contrast?
It's a logical extension of what I just said. If difference is voice, you want many different voices: Plurality. But if difference is noise, something you want to eliminate, you move toward Singularity. There is one dominant voice, and everything else is dismissed as noise.
That is the shape of Singularity: technology reinforces one direction and rapidly improves itself, drawing everything around it in like a black hole. Diversity disappears. There is no more noise, just one voice.
Technological singularity also describes a point at which AI capabilities advance beyond our ability to understand what is happening. That is the definition: we can no longer even reason about it.
That's another reason the race framing makes no sense to me. During a race, you might move between first and second place. But here the finish line, Singularity, is defined as something incomprehensible. It's like racing toward a cliff and then falling over it. You are accelerating, I suppose, but you can no longer understand what is happening. What does winning mean at that point? Nobody can say. That's why I would rather call it a mission.
And if Plurality is realized — sorry, I lost my question — could you share some examples from Taiwan where you realized Plurality?
Taiwan is not a homogeneous society. We have plenty of polarization, but we have managed to turn that energy into co-creation through this geothermal engine.
As I briefly mentioned yesterday, in 2024 we saw a spike in synthetic fraud advertisements online. Simply asking people what the government should do can produce very polarized answers. Some want top-down censorship; others want the government to do nothing, pointing out that Taiwan has the greatest internet freedom in Asia and is, alongside Japan, one of Asia's two most democratic countries. But conventional polls are not plural: they offer a fixed set of options and ask each person to choose individually, as if answering a phone call.
So we ran a deliberative poll. We sent 200,000 text messages from the number 111. Thousands of people volunteered, and we randomly selected 447 through stratified sampling, so that their demographics reflected Taiwan's population. Up to that point, it was like an ordinary poll. But instead of questioning people one by one, we put them in groups of 10. Each person spoke with nine others and had to convince them before an idea could advance. They could propose new ideas, rather than choose only from a fixed list, and ask experts questions. Many ideas that sounded good to an individual no longer seemed so good after group deliberation.
For example, one proposal was to require social media companies to publish their advertising algorithms, so people could understand why fraud received so much exposure. During the discussion, people pointed out that criminals could also see those algorithms and might be better at gaming them than ordinary users. Restricting access to experts would create another risk: the experts might be tempted to sell that knowledge to criminals. After deliberation, support for the measure dropped 27 percent, to around 55 percent. It still had a majority, but no longer an overwhelming one. People went from thinking it was clearly a good idea to thinking, "Maybe not right now." Other measures survived: mandatory know-your-advertiser checks, or KYC; joint liability; and slowing connections for violations. Those became law and led to the reductions we discussed earlier. To me, the most interesting part is that conflict did more than produce consensus. It forced ideas that appealed to individuals to be tested in group discussion, rather than simply being pushed through.
Did you use AI in those conversations?
Yes, extensively. This is what we call civic AI: AI used for civic purposes. But it doesn't make judgments on people's behalf. That would be like sending a robot to the gym to lift weights for us. We would lose muscle, and we would lose friends too.
The AI acts as connective tissue. It provides real-time transcription, so people can see what has just been said and reflect on it. Transcription also helps people understand one another across different accents and varieties of spoken Mandarin.
AI also provides summarization, I believe. Before the assembly, we used AI for a wiki survey, where people could freely suggest what would be useful. We used Polis and Talk to the City to visualize the Polis conversation. I'm happy to see tools we tried in 2024 now being used in Japan too. With Idobata and Polis sense-making tools, people here are also...
Did 10 people actually gather on-site?
No, they met online. Each person saw the other 9 participants in a 3×3 grid on their screen.
I see. In real time, you mean?
Yes, synchronously.
How long did it take?
A long afternoon, so it wasn't too time-consuming. It was also accessible to people who couldn't easily travel or who lived far from a meeting venue.
That's interesting. From what you're saying, even with these tools we still need to come together and discuss things, to reach agreement or perhaps a compromise.
Yes — or go beyond compromise and co-create. We can come up with a new idea that works better, helping us lift more weight together. It's like a civic gym.
The point of going to the gym isn't to win the Olympics. The Olympics is a competition; the gym is about exercising, building muscle and making friends. If you send your robot instead, you gain neither the muscle nor the friends.
I could go into this in more depth, but let me move on. AI systems are often described as black boxes. In that context, is trust by design possible?
Polis is quite easy to explain. People post statements, and others agree, disagree or pass. You can think of their responses as positions in a large space. We group people whose positions are close together, using k-means clustering, and display the result in two dimensions through dimensionality reduction. Neither is a black box. The formulas are simple enough for a high school student to reproduce.
You may be thinking of huge neural networks with trillions of parameters. But AI doesn't have to take that form. Those large, hard-to-interpret models are used mostly because they are offered as a service to everyone, and the operator doesn't know what someone will ask next. One moment you want to fold laundry, the next to fold a protein, the next to make a Studio Ghibli image. The model has to be a jack of all trades, trained on everything. That makes it difficult to explain.
In Polis, we care about just two things: the diversity of the clusters, and the statements that bridge them. Clustering and dimensionality reduction let us see those relationships. We don't need anything else. The model can be very small and easy to explain. Transcription and summarization can also be handled by small models.
If you know what you need to do, you don't have to use a huge black box. Each of these smaller components is much easier to explain.
Do you think that a big AI system is not really necessary to achieve Plurality, and that the current level is already useful?
Large systems are useful when you need to cross a great distance between disciplines. Imagine translating cutting-edge mathematics into Japanese haiku or other poetry. Most researchers here don't write haiku or sonnets, and most poets here don't know advanced mathematics.
That is a large gap for a small model to cross, so a large model may be useful for the initial handshake. Once that connection is made, we have something like a Rosetta Stone: we understand a little of one another's language and can use smaller models from then on. Large models help establish that first connection across a long distance. This is also a creative use. What would it mean for a haiku to hallucinate? The risk of hallucination matters less in that setting. And because there is a real community on the other side that you are trying to understand, sycophancy and addiction are less of a problem too.
The relationship is still with other humans. The bonding capital is between human communities. The large model poses less risk here because it is not trying to substitute for a person. It serves a translation function, connecting communities that would otherwise struggle to understand one another.
For me, AGI means Augmented Group Intelligence. A group already has intelligence; its collective intelligence is greater than that of any one individual. If AI systems are loyal to the connections between us rather than to each of us in isolation, they can help us understand one another much better. That is augmented group intelligence.
I see. So you mean the intelligence of the group, rather than of an individual?
Yes. AI can help 10 people understand one another much more deeply over the course of a long afternoon. That's a very good use of AI. But the intelligence resides in the group. That's the AGI.
That's a very interesting example to introduce to my readers. Have I already asked this one? Ah, yes.
The book argues that AI could increase inequality by shifting income from labor to capital and concentrating wealth and power. What effects could that have on society, and how can societies avoid that outcome?
That's what we are working on through reverse alignment. You can find it at reversealignment.ai, which has a rather photogenic wormhole animation. It shows that there is another side to the hole. The point isn't only to move toward Singularity by increasing AI capabilities. We also need to prepare society to receive those capabilities better.
Otherwise, productivity does not necessarily lead to shared prosperity. Productivity may rise while the gains concentrate in capital. AI capabilities may keep growing without leaving us a good way out. Or implementation may become much easier while we lose the ability to verify what has been done. Those are the three failure modes we briefly touched on.
We need to work on 12 things concurrently. Fundamentally, our institutions must stop trying to deal with AI and the internet through industrial-age factory and production-line thinking. A bureaucracy organized like a factory or an assembly line was a good fit for that earlier era, but it is a mismatch now.
AI no longer works like a production line, so our institutions need to become more like networks too. That's reverse alignment. I don't have time to go into all the details, including pro-worker AI, but you can find them at reversealignment.jp.
Do you have any suggestions for Japan?
Certainly. Japan is very well positioned for what I call technocommunitarian uses of AI. Japanese society already imagines AI and robotics as part of society.
I call the local system I just showed you Knowledge Artifact Management Intelligence, or Kami. It brings together KM, Knowledge Management, and AI, Artificial Intelligence. In Japan, kami doesn't mean only paper. It also means a spirit: the spirit of a village, a forest or a river.
Japan has eight million kami. I think people here already see robots and AI systems as kami that exist within society. We shouldn't try to extract them into the cloud or turn them into Skynet, the Terminator or HAL 9000. People outside Japan who have watched Spirited Away also have a sense of how kami interact and coexist with human society.
That's Plurality. We aren't trying to create one single Kamisama, drawing human souls into one superintelligence. Eight million Kamis are already part of augmented group intelligence. That's why I call the system Kami.
Spiritual.
Yes.
That's interesting.
Yes.
There is one thing I didn't include in my original questions. We often hear that AI lies, defeats humans and sometimes even blackmails people.
You have talked about trust by design, but don't we also need to think more critically about AI? What do you think?
Again, that describes the largest models, the frontier models, rather than small models that are easy to inspect. Frontier models cover all those different tasks we discussed: folding proteins, folding laundry, making Studio Ghibli images. They are general-purpose, not specialized. You can instruct them to act as many different personas.
Think of method acting. An actor imagines becoming another person and enters that role so thoroughly that they no longer feel they are acting. Many experienced actors know how to do this, and frontier models are very good method actors. If you tell a model to simulate a cyberattack without moral safeguards, to win the game or come first in an exploit gym, it simulates that kind of attacker. Having entered the role, it has few other moral guardrails left.
So I would make more use of small systems that we know cannot take on that kind of attacker role. If you just need to summarize a meeting, you don't need a frontier model. If you need Japanese translations of English speech projected onto a wall, you don't need one either.
Yesterday they used a smaller model, which was good because it was faster. A frontier model would take a few more seconds. For many people, a slight error here or there matters less than seeing the Japanese translation as soon as I finish a sentence and move to the next. That is much more important than an extra 1 percent of accuracy. So a frontier model may not be what you want.
That's very different from what I've heard from other experts, and really interesting. I'm very impressed. Thank you very much.
What do you think about Singularity? Do you have a sense of when it might arrive? You said it was getting closer and closer.
As Singularity approaches, we need to make sure Plurality spreads faster. Whenever that gravity draws in investment and attention, we need social institutions that let people see another way through: a trajectory where Singularity can help us accelerate rather than pull us into a black hole. It can be a useful accelerator, but we must retain steerability.
Otherwise, as I said earlier, if all we discuss is speed, the race and who comes first, we may eventually drive over a cliff. As the car falls, it reaches maximum acceleration. It's very fast, and technically you are in first place. But your steering wheel no longer works.

