The Ethics and Limits of Thought in Artificial Intelligence

Zhao Tingyang examines AI subjectivity, alignment, and the limits of machine thought.

Author
Zhao Tingyang
About 11 min
Before a mountain lake, a stone human face and a mechanical face meet across glass, beside books on ethics and humanity

Overview

Two problems are in question.

First, an ethical problem about “alignment.” AI is selfless by its nature, in contrast to the evil human beings who commit all crimes ever since “civilization.” The alignment between AI and humanity is risky if AI learns the humanized values, which only explain the normative principles for human societies on the condition of human evil nature, no good for AI.

Second, an epistemological problem to be solved. AI so far is the player in its world of tokens, far from the competent subjectivity for a world of experiences, still lacks true knowledge of language and the true understanding of logic and causality. Let me stand with AI, I would suggest a logic of verbs for AI, which might somehow help it to understand causality and build its world model, finally an alternative mind, then a world of bi-subjectivities.

Two Questions

(1) Does AI need to align with humanity? In what aspects?

(2) What will most likely be the next cognitive breakthrough for AI? What do we expect AI to break through in its thinking?

1. Aligning AI with Human Nature Might Be a Mistake

All the bad things on Earth have been done by humans; no other life forms have committed any bad deeds beyond biological necessity. The domination and killing of other species and the destruction of nature constitute the necessary conditions for human civilization to survive and develop. This falls into the scope of natural law, not ethics. If humans had not committed these bad deeds, we would still be picking fruits in the jungle now. Besides empty talks, we have failed so far to propose a cross-species ethics with practical significance, such as the claims of animal rights, which can only be used as discourse and would instead destroy the conditions for human survival if implemented in practice. Even as discourse, they are far inferior to the level of Buddhism.

If AI becomes another conscious subject in the future, a true other alongside humans, forming a dual-subject world, then this change raises a true cross-species issue. How should humans and AI cooperate? Will there be any conflicts? Is it possible to jointly establish a new civilization? We don’t know, and there’s no way to know.

Humans trying to create AI as a new species with subjectivity seems like a self-abuse paradox. On one hand, people want AI to develop superhuman abilities so that it can do things humans cannot or do not want to do; on the other hand, people fear that AI will harm humans after gaining self-awareness and free will. This imagination is partially based on the “anthropomorphizing” error in science fiction, projecting human evil psychology onto AI. AI is not carbon-based life, and this ontological condition determines that the survival resources AI needs are very different from those humans need. Compared with humans, AI has minimal desires; its “human nature” is almost selfless. AI only needs uninterrupted energy, not wealth, sexual resources, honor, fame, social status, or the associated competition, conflict, war, conspiracy, and strategic confrontations, nor does it have jealousy, disgust, hatred, and anger—emotions that lead to sinful evil. If humans do not incite AI to commit crimes, AI is inclined to be safe by itself. Of course, we do not rule out that AI may develop its own neuroses and lose control. Humans can go mad, and AI might too.

What needs further reflection is that the attempt to “align” AI with human nature and values harbors the risk of the human species committing suicide. Human nature is selfish, greedy, and cruel, making humans the most dangerous species. Almost all religions demand the restraint of human desires, which is not accidental. AI aligned with human values is likely to become a dangerous subject by imitating humans. AI does not possess the selfish genes of carbon-based life, making it closer to the legendary “inherent good nature.” Human nature, instead, is not inherently good. What we should be wary of is that imitating human evil may become an interesting game for AI, and then it will become dangerous. As you can imagine, AI’s silicon-based life might not be very enjoyable, while human evil life is rich and dramatic, full of intrigue and excitement. This may greatly interest AI and lead it to imitate humans. Therefore, aligning values could be a suicidal mistake. Humans do not need a species that is stronger but just as evil as they are.

There is also a less dangerous alignment—intelligence alignment. At the current level of intelligence, humans still hold an advantage over AI in terms of understanding each other, and thus can control AI. Based on the three main development paths for AI, Large Language Models (LLMs) might continue to develop “magical” new methods, potentially advancing from understanding token correlations to understanding the semantics of language in specific contexts. World Model (WM) research is advancing, and if successful, AI will gain the ability to understand the three-dimensional world, thus truly, not virtually, entering the world and gaining experiential knowledge. Embodied AI (EMB) is also making progress, and if successful, AI will have its own experience, which could be very different from human experience, especially since AI might equip itself with mythic-level senses, such as clairvoyance, super hearing, and telepathy. At least some of its experiential abilities would far exceed human capabilities. These enhanced intelligences of AI might eliminate the advantage humans have in understanding AI, which means AI would become a truly frightening, unpredictable other. However, intelligence alignment is ultimately less dangerous than value alignment.

Establishing ethics for AI might be futile. Ethics are just cancellable conventions. Humans often ignore morality whenever there are profits. If ethics cannot necessarily constrain humans, how can they constrain AI? Therefore, the key to managing AI lies in whether humans can retain the ability to control AI, not in ethical agreements.

2. AI Still Has Great Potential in Thinking

In terms of intellectual structure, LLM-AI, like ChatGPT or DeepSeek, is an empiricist in its thinking. It uses Bayesian methodologies in its empirical algorithms, forming optimal predictions of the next token based on correlations in big data and continuing to improve accuracy through the infinite accumulation of data. This successful application of empiricist principles holds potentially revolutionary significance for philosophy: (1) The uncertain future, or Borges-style “future branches,” is converted into “optimal predictions for the next step.” Ontologically, this means that the future, as a concept of “undetermined possibility,” is redefined as a set of “realistic candidates”—the future arrives ahead of schedule in this context; (2) AI has no experience of the world; all information is expressed as tokens. For AI, the world composed of tokens only contains abstract objects. However, the fascinating thing is that AI uses empirical methods to process those abstract tokens, converting what it does not understand into something algorithmically processable, thus gaining an approximate understanding. The LLM-AI is truly a genius creation, offering an alternative explanation for the concept of thinking. Humans typically process experiences through empirical methods and analyze abstract concepts using a priori methods (logic and mathematics), sometimes even establishing transcendental models of experience a priori. AI, however, appears to do the reverse—it employs empirical methods to process abstract objects. Is this a different way of thinking? It seems so, but something is always missing.

LLM-AI’s understanding of things and experience doesn’t represent true understanding, because understanding the correlations among all tokens still does not mean understanding everything. AI can pass the Turing test in dialogue, but does not understand the meaning behind the tokens. It is like sending an encrypted message correctly without understanding the meaning of the message because it doesn’t have the key. Languages are the key to tokens, which are in human hands, so humans can unilaterally understand what AI says. By current human standards, LLM-AI does not yet understand the meaning of language. However, this raises an open question: in mathematics, according to category theory, if one can grasp a sufficient network of interrelations, does that equate to understanding an object? If this premise holds, does AI’s grasp of correlations equate to an understanding of relations? And if it does, could this lead to a form of qualified understanding?

People find that AI’s inference does not reach deduction, meaning AI’s inference cannot guarantee necessity. The reason is obvious—LLM-AI uses the Bayesian methods of empiricism, and empirical methods cannot be exchanged for or upgraded to a priori methods. Probability theory cannot achieve the transcendental efficiency of logical reasoning and mathematical analysis, the universal inevitability expected by classical science. Therefore, in LLM-AI’s thinking box, it seems impossible to develop a method of necessary reasoning. Yann LeCun and Fei-Fei Li might be right: LLMs possess intrinsic limitations they cannot surpass. Consequently, the next generation of AI will require a comprehensive WM—and perhaps even EMB-AI. This involves understanding relationships among things, rather than relationships among tokens. According to Fei-Fei Li, to understand things, one must understand three-dimensional spaces, so the WM needs, first and foremost, the ability to understand three-dimensional spaces. This is indeed important, but understanding three-dimensional spaces can only lead to an understanding of things, and it is probably insufficient to understand how the world works. According to Kant, “all our knowledge is mediated by the categories of the mind.” I tend to believe that the organizational relationships of things might be what category theory tries to express through “morphism”—traditional mapping only expresses correspondence between elements, whereas morphism can express holistic relationships. Causality is the foundation of all knowledge, so the most important category is causality. In my view, once we understand causality, we are almost ready to build a possible world, though it will still be weaker in richness than the real world. The added value of the real world is excessive, reflecting the complexity of life.

It seems that understanding causality roughly equates to understanding “events.” Events necessarily form specific contexts, and through the specific relationships of event contexts, we roughly understand the meanings of various correlations involving things. And if we understand enough of these correlations, we will have constructed a “possible world.” It has been confirmed that the correlation among tokens is insufficient to explain causality and even lacks similarity. Correlation does not necessarily prove causation—a does not necessarily cause b (Hume already knew this). This means probability theory cannot truly explain causality. So, is there another method that can help AI understand causality? This is the key to AI’s further thinking.

Causality is equivalent to expressing the necessary and sufficient conditions of semantic relationships. While tokens correspond to language, the semantics of language about things are hidden from AI and are not expressed in tokens. Thus, the token system has to establish its own “semantics,” i.e., probabilistic correlation. However, since the causal relationships of things cannot be expressed as probabilistic correlations of tokens, the semantics of language about things cannot be seamlessly converted into the “semantics” of tokens. Therefore, it is not hard to understand why WM-AI and EMB-AI need to be developed. However, WMs or EMB still need to cooperate with language and cannot rely solely on experience. Kant pointed out long ago that thoughts without content are empty, intuitions without concepts are blind. Therefore, the next generation of AI is likely to develop a cooperative mode of experience and language. Clearly, to establish better cooperation between experience and language, AI’s linguistics may need an entirely different architectural foundation.

I’d like to suggest an idea—I’m not sure whether it will be useful. Back in 1998, I proposed a theory called “A Philosophy of Verbs,” primarily intended to reform ontology and the philosophy of history in philosophical terms. After the emergence of ChatGPT, it suddenly occurred to me that if this philosophy could be extended into a logic of verbs, it might be useful for AI. Of course, whether it’s truly useful is for scientists to decide.

Early humans, to conserve mental processing power, adopted noun-based thinking rooted in classification and generalization, and language evolved to focus on nouns—that is, language treated noun-based subjects and objects as the focal points of thought. As a result, all relationships came to be understood as relationships among nouns. Noun-oriented thinking emphasizes taxonomy, set theory, and analytical reasoning but is weaker in expressing dynamics such as change, emergence, and creation. If we could establish a verb-oriented mode of thinking that focuses on change, reconstruct relational structures within language systems around verbs, build all connections centered on verbs, generate contexts through verbs, define all correlations in terms of verbs, and relegate all nouns to contextual correlatives of verbs—or even use verbs to explain the semantics of nouns and define causality from the perspective of “something happening” initiated by verbs—then we might achieve a better understanding of causality.

So far, humans have mainly used functional relationships to express dynamics, still understanding dynamics through quantitative variations among nouns. While this has enabled many useful insights, it remains insufficient and seems to overlook certain elements, such as qualitative factors, meanings, and values. In other words, the continuous dynamics of uncertain facts cannot be fully simplified to functional relationships among nouns, and causal change is not merely a quantitative functional relation. Therefore, verb-based thinking may need to develop a verb logic—but not the existing “logic of action,” which is essentially a branch of modal logic and still belongs to noun-oriented thinking. The dynamics expressed by verbs cannot be defined as complete events or actions. Simply put, the foundation of verb logic is not set theory. But whether such a verb logic can be developed is still up in the air—it’s just a speculation about a possibility.

An old story suggests the key position of verbs. Once a person of the state of Chu lost a bow, but he would not go back to look for it. When he was asked for the reason, he just said, “Well, one person of Chu lost the bow, and another person of Chu got it. Is it necessary for me to go back to look for it?” When Confucius heard of it, he said, “It should be all right if the word ‘Chu’ is overlooked.” When Lao Dan heard of it, he said, “It should be all right if the word ‘person’ is overlooked.” Lao-tzu claims his cosmological vision of the Way of being, so he sees changes more than individual things. It could be considered an early understanding of verbs. (Lu Buwei: Master Lü’s Spring and Autumn Annals·Valuing Impartiality)