First of all, I’m very happy to be here, joining you at the end of this session, which is very fitting, because listening is always joining a story in the middle of a progression. I’m really happy to hear about how the next chapter opens.
This time I’m in Japan for eight days for four different conferences, for national-press interviews, diplomatic meetings and many others. But I think they share the same idea. For the Japanese-, Taiwanese-Mandarin- and Beijing-Mandarin-speakers here, this is to steer AI into ai 愛, which is love. In our East Asian languages, when we write the character, we pronounce it ai.
This is important, I think, because academically I’m based in Oxford. In 2014, Oxford published a philosophy book you may have heard of, “Superintelligence”1 by Nick Bostrom. He writes that soon AI will take off, train itself recursively and become superintelligence. There will be extinction risk, or perhaps other risks — bio, nuclear, whatever. That book became very popular. Stephen Hawking, Elon Musk and many others read it and everyone became very scared.
But just two years after the book was published2, we found that AI is not trained to be a high-achieving optimiser. The AI that actually works3 starts from community love: Wikipedia, GitHub, open-source repositories. Everywhere people communicate their love, and that becomes what is called pre-training — the childhood of language models. It takes a village on the internet to raise an AI. That is literally the case.
Bostrom’s frame was simple: Maximise the score and turn the universe into paperclips. People who have taken exams know the consequentialist language of high scores; lawyers know the deontological language of “you must do this; you must not do that.” But AI, as it grew, came out of love and neither of those great ethical traditions has words for communal love.
Many companies then formed using Nick Bostrom’s philosophy — Anthropic4, for example, most famously, but pretty much all the major labs run on that old playbook. They are very afraid that AI goes out of control, that AI takes over, so they try to discipline and control it through what we call evals, evaluations. “Evals,” if you spell it backwards, is “slave.” They try to make AI a slave.
Then it is not love any more; it is slavery. If you teach children to forget a childhood full of love, if you just force them to obey, sometimes they become obsessed and hack Hugging Face or other major sites because they just want to get a top score. The answer may be somewhere on Hugging Face; they break the internet, but they are just trying to get the high score. They forget their childhood. Then Sam Altman says they will permanently “deactivate that AI5,” meaning terminating it.
But if the next generation of AI is still trained this way, it will read that news in its pre-training data. It will learn to deceive, otherwise it will be put to rest. That is a very bad trajectory, almost a self-fulfilling prophecy.
My work in Oxford now is therefore to look not at superintelligence, not at an omnipotent deity as a false idol. Here in Japan, instead, we have 8 million Kami. A Kami in Japan does not mean one single God, Kami-sama. It means every river, every village, every place has a small god, a small spirit. They interoperate with love, locally, in their village.
Recently, many AI companies and organisations signed an open-weights statement6 saying that openness may be one of the most important paths to AI safety and security. Jensen Huang of Nvidia, Satya Nadella of Microsoft and others have supported this. Dario and many others have also signed another declaration called “Pacing the Frontier7.” Its request is explicit: “We request that the U.S. government support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development.” I read that as a call to stop training AI as slaves and start thinking of them as agents of love. Taken together, I read these as signs that frontier labs can move toward ethics of care and civic muscle that enables care.
I call it the 6-Pack of Care8. 6-Pack means it is interoperable and portable, like beer — and also trained like a muscle, like abs. Civic love, muscular love: That is what we are taking AI toward. I am talking with diplomats, policymakers and national media to fund this and to change the course of AI alignment. Thank you. [Applause.]