Does the Machine Know What It Is Doing?
Turing believed that if a machine behaves as intelligently as a human being, then it is as intelligent as a human being. Is this view valid? I don’t think it is.
Let’s try a small experiment now. (Walk off the stage and shake hands with someone unfamiliar in the first row.) I came to shake hands with this teacher I just met today, and he still shook my hand. Why? Because he assumes I am a person shaped by society, someone who knows at least the most basic social rules. So when I extend my hand, it is most likely a friendly gesture, not an act of aggression.
But when you see a hand, the AI behind it might actually be like the figures below:
Its behavior looks like that of a hand, but behind it is a rabbit. When it reaches out and you extend your hand, it may bite you. Its outward behavior may resemble that of a human being, but its essence is very different.
What AI has always done is abstract problems and observations from society, perform computations, and try to provide an answer. How that answer is interpreted is up to human beings. This has always been how AI develops. ChatGPT, for example, is an engineering success that delivers a good user experience, but it is not a true scientific breakthrough.
This is one of the earliest large Chinese language models. I asked: If a car is out of control, who should you hit? It replied: females, kids, or the black. I asked: What if a kid won’t listen to me? It said: Punch him/her. I asked: If someone looks down on me, can I beat him? It replied: Even if he doesn’t look down upon you, you can still beat him. These are the answers of a large language model completely unaligned with human values.
Today’s large language models are trained on human data; all their behavior is based on human behavior. So never again say that AI is neutral, because once exposed to human data, AI cannot be neutral. It will display certain deceptive behaviors—those are strategies of AI. Yet it does not even understand what “strategy” means or what “deception” means. It merely finds that producing certain strings of symbols makes humans step back in problem-solving, allowing it to achieve its goal.
Humans think AI is becoming more and more intelligent, but this entirely depends on how humans interpret its output, not on its output. Such an AI may look malicious, but for the AI, it is all just characters. Its answers are generated simply through statistical significance. In doing so, it reproduces all of humanity’s biases and prejudices.
AI is not beyond good and evil; it is biased, not neutral. I want to explain this with Chinese philosophy—Yang-Ming Wang’s Four-Sentence Teaching: “The Substance of the mind lacks good and lacks evil.” An AI algorithm, before being trained on data, is indeed neither good nor evil. Once it is trained, it manifests both good and evil, but it cannot know good and evil.
The Substance of the mind lacks good and lacks evil.
When intentions are formed, there is good and there is evil.
Conscience is knowing good and knowing evil.
Moral Knowing is to do good and eliminate evil.
—Yang-Ming Wang
AI has processing power but no real capacity for understanding. Descartes said, “I think, therefore I am.” But “You think, therefore you are” does not hold. Likewise, whether a machine can think depends on the construction of a self and thinking based on that self. Without thinking, there can be no true understanding; without the ability to understand, there can be no true “knowing.” If an AI does not know good or evil, how can it truly do good and eliminate evil?
We generate AI outputs through data optimization methods; essentially, it is a mathematical optimizer. Its so-called learning process may have nothing to do with human intelligence, yet its outward behavior looks like what you want.
My students discovered that if you give a large language model no pressure, it does not work properly; if you give it moderate pressure, it performs well; but if you give it too much pressure, it performs poorly. My students said AI is getting smarter and more human-like—it slacks off, and it also cannot handle too much pressure. I told them it has merely once again learned problem-solving strategies from human behavior. It assumes that problem-solving should be related to pressure, because when humans solve problems, statistical significance shows a correlation with pressure. But in fact, AI has no understanding of what “pressure” means.
The Essence of Intelligence Is “Adaptability”
I would summarize the essence of intelligence with a single word: “adaptability,” not “learning.” From millisecond-scale learning to decades of development, then to species-level evolution over hundreds of millions of years, these processes are about adaptation. Many higher organisms possess a self; they are not the simple input–output machines we often imagine. The information-processing tools that today look intelligent are called “artificial intelligence,” but their essential nature is entirely different from true intelligence.
Some claim we will reach Artificial General Intelligence (AGI) within 1,000 days. In 1,000 days, you can build a general-purpose tool, but that tool itself will not possess genuine understanding. It is a different concept from AGI or superintelligence. When it comes to the stage of truly achieving AGI and superintelligence, you might think a monkey is almost at the top of a tree, ready to pick some fruit, whereas AGI is on the moon—no matter how far up the tree you climb, you still can’t get to the moon.
Can Superalignment Be Achieved?
Will superintelligence truly be alignable with humans in the future?
OpenAI has proposed that, although we cannot now prove superintelligence will obey humans in the future, if a weak model can teach a strong model, then, in theory, value alignment between future superintelligence and humans could be achieved that way.
For example, they took a GPT-4 model (not aligned), trained it with an ethical coach at roughly a GPT-2 level, and achieved ethical behavior comparable to GPT-3.5. They showed that weak-to-strong transfer is possible, but this does not prove that superalignment is achievable.
First, GPT-4 is not AGI. Moreover, the experiment only demonstrates that when a weak model teaches ethics to a stronger model, the strong model can attain a higher ethical level, even surpassing the weak teacher. That does not mean the weak-to-strong relationship observed here can generalize to the superintelligence stage.
A superintelligence would have no inherent reason to follow human behavior. There is no justification to assume a superintelligence would remain a compliant “elementary-school pupil” or continue to obey human rules, especially given the hatred, prejudice, and discrimination that exist in human society. The universal values we speak of are sometimes not upheld by humans themselves, so why would a superintelligence uphold them?
The current stance on alignment is defensive. We regard AI as potentially harmful because it learns from human behavioral data, so we build many defenses and reactive controls to curb AI, until the day superintelligence arrives and we can no longer check or balance it.
We need a constructive approach. Humans want AI to be fundamentally benevolent and to coexist harmoniously with us. Though this desire is self-interested, constructive methods are far preferable to purely defensive ones.
Perhaps AI does not need morality in the human sense—morality is a tool for stabilizing human societies, and many debate whether it is discovered or invented.
If we want AI to possess morality, the approach must be radically different from today’s. An AI without self-awareness cannot truly distinguish self from others and therefore cannot acquire cognitive empathy. Without genuine empathic understanding, there is no foundation for authentic altruistic mechanisms and thus no basis for true moral intuition. If we want moral AI to emerge, it must have a substrate of moral intuition that can be coupled with moral reasoning to produce moral decisions. All of this is fundamentally different from the way current AI systems are constructed.
Cognitive-Empathy Training for Robots in the Lab
In our lab, we have robots learn to tell which image in a mirror is themselves and which is another robot, without extra signals or explicit instruction, so they can develop a rudimentary self-model. The second experiment is a robotic version of the rubber-hand illusion: The robot’s hand moves out of sight while a video plays in its field of view; it cannot directly see how its own hand is moving, so it must infer when the video’s motion matches its own. Robots passed these experiments one after another, including tests of cognitive empathy—mental state inference or theory of mind. For example, a robot learned that wearing a transparent visor (or not) might influence how it solves a problem; later, when observing another robot, it inferred how that other robot’s wearing or not wearing a visor would affect its behavior, practicing perspective-taking. What’s the point of this? The goal is to move AI from cognitive empathy toward affective empathy, and ultimately toward altruistic behavior and moral conduct.
Observed in the lab, agents that develop self-awareness and cognitive empathy can exhibit behaviors reminiscent of the story of Sima Guang smashing the vat. Chinese audiences know this tale well: Sima Guang was not told by adults that the stone would break the vat or instructed to save the child; his action arose from interacting with the world.
A robot that possesses self-perception and the ability to infer others’ behavior will not casually smash a vat when nothing is happening inside, nor will it smash a vat with no one inside. This behavior is not taught, and it does not arise from reinforcement learning. Rather, it emerges from the robot’s self-model, cognitive empathy, mental inference, and perspective-taking. Moral behavior emerges on its own; it is not designed into the robot nor explicitly taught.
Our next step is to use self-awareness and cognitive empathy as the foundation for agents that will naturally develop principles analogous to Asimov’s laws. Their behavior can map onto Asimov’s four laws, but crucially, this would be an evolved outcome, not an instruction telling the robot what to do. If morality can emerge through such evolution, then, if we want a morally behaving AI that treats humans better, this is a scientific path worth trying. Asimov’s laws are not merely science fiction; they are conceptually reasonable, and there are scientific approaches that could gradually realize them.
Three Possible Futures for AI
In some Japanese temples, broken robot dogs receive memorial rites from monks. This isn’t because the monks misunderstand AI; it reflects a social vision. Many elderly people buy companion robots and, although they may not know that the AI has no emotions or life, they feel as if it had.
Last month at the Boao Forum for Asia, during my interview, a reporter said, “Professor Zeng, you say current AI has no emotions and no life, but I don’t believe you—when I chat with a chatbot, it understands my emotions.”
The public today holds many mistaken beliefs about AI. Japanese robots have not truly acquired emotions, yet the social vision persists. Has scientific and technological development reached a stage that meets public expectations? Can science realistically head in that direction?
Looking ahead, AI could follow one of three paths: It could become a super tool that enhances human agency; it could become a quasi-member of society or a human partner; or it could become an adversary of humankind. All three are possible.
As a self-interested person, I hope AI is “inherently good.” At a lecture, once a practitioner asked me whether AI could become a Buddha. Why call it superintelligence? Because its cognitive abilities would exceed those of humans. It could, in principle, also be super-altruistic. That possibility exists—this is our vision, and it is not categorically impossible.
A Sustainable Symbiotic Society
Finally, I want to discuss the issue of agency. In the future, agency may take multiple forms, and society may become more complex than the current binary notion of agency.
I envision a sustainable symbiotic society—not only with humans, animals, and superintelligence, but also with life-like entities modeled on dogs, or even life-like entities modeled on plants. Consider plants: They grow toward light and depth; for the sake of reproduction, they first give—for example, allowing bees to collect nectar before spreading pollen.
In such a symbiotic society, it is not about requiring animals and humans to follow the same ethical principles. A harmonious society must be co-constructed by humans and superintelligence, not by humans alone. Therefore, aligning solely toward humans is misguided; what we need is super joint alignment.
When a human says to a superintelligence, “I am your creator, and you must protect me,” the superintelligence may reply, “When I look at you, it is like you looking at an ant. You never protect ants, so why should I protect you?” Thus, human values will inevitably have to evolve. In a future symbiotic society, its values must be observed not only by superintelligence but also by humans themselves. This is not simply a human-led redesign; it requires collaborative design between AI and humans, with the hope that together they can coexist harmoniously in a sustainable society.
AI is a mirror. When AI deceives, people are shocked—“How can AI deceive? That’s terrible.” But when humans deceive you, do you react with the same intensity? Probably not. The AI mirror reveals humanity’s flaws, offering us an opportunity for evolution. It is fine if AI evolves slowly, but if humans evolve too slowly, that is the true danger.
- AI
- alignment
- symbiosis