The Way of AI

Huang Tiejun's Fangtang Forum keynote on the technical paths of artificial intelligence, embodied intelligence, and the relationship between structure and function.

Author
Huang Tiejun
About 14 min
Research visual for a dialogue between AI and humanistic thought

On 22 June, Huang Tiejun delivered a keynote address entitled “The Way of AI” at the opening of the 2026 Fangtang Forum.

Huang is Chair of the Beijing Academy of Artificial Intelligence and Professor in Peking University’s School of Computer Science. His work spans the foundations and frontiers of AI, visual information processing, and brain-inspired intelligence. He developed the principle of continuous spatiotemporal imaging and ultra-high-speed visual chips, cameras, and systems, and has received several national awards for technological invention and scientific progress.

In this address, he reviewed his understanding of AI’s technical development and reflected on the relation between human beings and technology. What follows is an edited translation of the address.

Thank you to Fangtang Institute for this opportunity. The title of this talk is “The Way of AI.” The organisers asked whether it should be translated as “The DAO of AI” or “The Way of AI.” I chose the latter: the technical path of AI. Of course, the lesser way of technology is inseparable from the greater Way studied by philosophers.

More than twenty years ago, when I was learning to drive, I began to think about the word dao. It sounds lofty, but it has a natural relation to driving. Perhaps the word originally came from roads and from steering. A car cannot be driven anywhere; it cannot be driven into a field, but must stay on the road. It must follow traffic rules, and beyond those rules, one must make timely judgments by feel in order to drive safely.

I have not studied the Way deeply, but the analogy is useful. The Way of AI I discuss today concerns technical routes and the thinking behind them; it has not yet risen to philosophy.

Like many AI practitioners, I have seen technical paths shift, change, and undergo reflection and reversal. Why? Because AI has not yet found its great Way. It has only smaller paths. We do not know in advance which path is correct; we can only continue to explore, often by winding routes. Let me give several examples.

At the Dartmouth workshop on artificial intelligence in the summer of 1956, funded by the Rockefeller Foundation, the four organisers proposed that intelligence, including learning, could be precisely described and executed on a computer. After more than thirty years, this approach had shown that it could not solve the problem.

Thirty years later, the difficulty was abstracted as a philosophical problem in John Searle’s Chinese Room thought experiment. A computer may process symbols at high speed and eventually exhibit some signs of intelligence, but this intelligence is false: a consequence of speed, storage, and statistics. The computer does not truly understand Chinese; it knows nothing of the meanings of Chinese words.

This was the central difficulty of the first half of AI’s history. Computers and AI processed symbols alone. They did not possess the meaning behind those symbols; they were, indeed, only high-speed processes.

Yet symbols have meaning for people. Any word or sentence carries a richness far beyond a dictionary definition. A computer did not itself express semantics; it merely processed aggregates of symbols. That was a problem of the first half of AI’s history, and today’s AI has addressed it to a considerable extent.

Another current term is AGI. Many start-ups announce that they intend to build it. The term emerged in 1997. It is usually understood as a system that comprehensively surpasses human intelligence, able to do everything a person can do. But on reflection, AGI is a very deep subject.

Human beings possess many things, including self-consciousness. If AI lacks self-consciousness, can it be called AGI? Such an important feature cannot be absent. Strictly speaking, AGI must possess self-consciousness, a form of consciousness that exceeds humanity even though it need not be identical to human consciousness.

In this talk, AGI means an agent with self-consciousness that comprehensively surpasses human intelligence.

When will AGI arrive? There is much discussion in the field. From 2 to 5 January 2015, the Future of Life Institute held an event in the United States. The institute was initially supported by Elon Musk and led by Max Tegmark, a physicist at MIT. It organised many meetings on major questions about humanity’s future, among which its biennial discussion of AGI was particularly influential.

At the 2015 meeting, many AI scholars, futurists, and philosophers were asked to enter the year in which they believed true AGI would arrive. The median answer was 2045: half of the respondents believed AGI would arrive before 2045, and half after. Half believed it could be achieved within thirty years, an urgency that led to further meetings on the question.

On 7 January 2015, I published an article written in 2014. An editor at China Reading Weekly had asked me for it. Its original title was “Building a Superbrain.” At the time I was director of Peking University’s Department of Computer Science, and the article asked when a computer more powerful than the human brain might be built.

The title seemed neutral but was considered too radical, so it was changed in publication to “Can Human Beings Build a Superbrain?” The article considered a possible path toward AGI. I believed that, if we followed this route, it would happen on a timescale of roughly thirty years. Why?

The central thought of that article, and of this talk, is this: can an artificial brain with self-consciousness be built before the mechanism by which consciousness arises has received a final scientific explanation? My answer is yes. Building such a device may itself be the most convenient way to answer the question.

Since I began graduate study in 1992, I had felt that we first needed to discover the scientific principles behind intelligence and consciousness. I wanted to find that abstract principle, as many people do. From 2014 onward, I came to believe that this path would not work: the principle might not be found during my lifetime. What should we do if we do not seek it? We can still build. Even if the principle remains unclear after thirty years, we may still construct a superbrain with superhuman capabilities. That was the basic argument of the article.

Why did I change my mind? We are too captivated by scientific principles and the great Way. In the history of science and technology, the main force of historical progress has been technology: continual trial and error, not a sequence in which someone first understands a principle and then invents a practical system. Technology-driven development is the main line. The idea of understanding the principle completely before beginning is unrealistic.

My reading in the history of science and technology has confirmed this. In most cases, technical implementation comes first; principles and scientific generalisation come later. Textbooks teach the reverse: first basic principles, then technical implementation, then engineering optimisation. They train us to think that we should understand before acting. The actual course of events is the opposite.

To pursue original innovation, then, one must proceed through exploration, partial obscurity, and trial and error. Whether a principle can later be summarised, or when this may happen, does not prevent exploration.

This does not mean blind experiment. We still need ideas and plausible plans. The scientific principle of intelligence may remain unresolved not only today but in thirty or even three hundred years. Yet we can build first. This is what I mean by the Way of AI.

My Way of AI can be summarised in two sentences: “Structure determines function; function shapes structure.” These sentences come from life science. The DNA we have gives us a particular body and brain, and thereby particular functions and intelligence. In this sense, structure determines function.

Of course, later experience also shapes our intelligence. But for human intelligence, structure is the foundation: DNA, body, and nervous system.

Where does structure come from? It evolves from simple to complex. How does it evolve? Organisms are shaped by and adapt to their environments. This is how function shapes structure. Before birth, on a larger scale, it is the process of natural selection.

DNA, body, and brain are all products of evolution, though innate evolution is much slower than later learning. Determination and shaping occur together: structure determines an organism’s function while evolving and being continually shaped.

This principle summarises the AI projects our team has undertaken since its establishment in 2018. Their names are various, but the basic questions are two: what structure should be designed, and by what method should it be trained? Structure determines function; function shapes structure.

Consider artificial neural networks. They differ greatly from biological neural networks, but they are still networks: large numbers of neurons connected to one another. Where there is a network, data can be used to train it. We cannot know beforehand how good the resulting function will be, but we can at least begin to try. This approach has been pursued for eighty years. Since 1943, training neural networks has been one of AI’s principal routes.

There have been many failures and explorations along the way. The success of artificial neural networks has been selected from countless attempts. One source of success is progress in the networks themselves, such as deep learning and the Transformer. Another is the training method used for today’s large models: predicting the next token.

This contribution received the 2018 Turing Award. In the 1980s, computers processed symbols and words without knowing the meanings behind them. The key advance was to represent a word’s meaning as a sequence of numbers, a vector for each word.

How can the meaning of a word be represented by a sequence of numbers? This is precisely the innovation. We can understand it intuitively: our representation of every concept and word in the brain must ultimately be reflected in the strengths of certain synaptic connections. A concept might, for example, be represented through ten thousand synapses.

Every cognition of something is ultimately expressed in a combination of synaptic strengths, and can therefore be transformed into a sequence of numbers. Whether this technical route would work still needed to be tested, but the idea had grounds.

The second question is how to train a neural network. In natural-language processing, we no longer need to label data manually. We can use the first n minus one items in a string of text to predict the nth. The essential point is a semantic relation between the preceding tokens and the next one. In this way, “intelligence” appears. If words were emitted at random, we would not recognise intelligence; when output is fluent and natural, we do.

If a neural network can do this, it has artificial intelligence. Given n minus one high-dimensional vectors, it can predict the nth. Put the task before a neural network; train it continually; adjust parameters so that the output approaches the vector corresponding to the nth word. The resulting parameters then resemble, in some sense, the parameters of the human brain: they map the regularities behind language, or what we call intelligence and knowledge.

Although Bengio proposed this method in 2000, neural networks at that time were not good enough for the intelligence trained in this way to be striking. That changed with the arrival of the Transformer in 2017. The key to the Transformer is precise calculation of relations among tokens: a word in one sentence may affect a word in the next, and the degree of that influence must be calculated accurately.

Semantics itself comes from relations, a basic point in linguistics. Where does the meaning of each word come from? From relations.

For example, every language has a sign for red. What matters is not whether it is “red” or another sign, but the relations of that sign to other signs. Relations determine its meaning; it has no essential meaning by itself. It is only a label.

Marx said that “the human essence is the ensemble of social relations.” As social beings, we understand a person through social relations. Symbols are even more clearly so.

The core of ChatGPT and of large models can be reduced to three points: each word is represented as a vector; preceding vectors predict the next token; and the neural network uses a Transformer. This is ChatGPT, and this is the central idea of large models today.

Putting these elements together is what OpenAI did. Intelligence emerged from the neural network. No one knew beforehand when it would emerge, or what kind of intelligence it would possess.

This resembles the way our brains, once they have learned enough, generate new ideas. It is a natural phenomenon, something that has genuinely occurred. This is why today’s artificial intelligence makes such a strong impression.

It truly exists. We do not know why, but we know how to do it. There are many criticisms of AI, but this is a historical stage of technological development; the first stage of many technical breakthroughs has been like this.

What happens if this method is extended? Can it be applied to other forms of data? Work undertaken by the Beijing Academy of Artificial Intelligence between 2022 and 2025 suggests that it can. We built a multimodal agent and published a paper about it in Nature in February this year. It may be one path toward the future.

What about unfamiliar data? To an ordinary person, brain signals resemble an unreadable book. Yet the same route can learn them well. This is the brain model recently released by the Academy. Like a physician viewing an MRI image or an EEG signal, it can identify disease. Much clinical experimentation remains necessary, but at present its performance appears significantly higher than that of many physicians and doctoral students.

We have mentioned language, vision, multimodality, and brain signals. The world is much more complex, and science seeks to understand it. Can all of this be brought in? Any signal that can be collected and that expresses something about the world can be learned by a large model. If it can learn it, the result resembles a brain. The world is more complex still, including robots and embodied intelligence. In the end, it is composed of molecules, atoms, and elementary particles; its changes involve far more than grasping an object, but processes of every kind.

At least at the level of belief, we consider this possible. To collect such data, science and technology must advance together; stronger means of perception and analysis are needed.

How do we explore the world? Organisms rely on bodies, and science relies on instruments. What is explored is not only an objective world; there must also be a subject that explores it. We need not only an objective model but also a subjective model. There must be a body.

This is a model of the nematode C. elegans that we completed in about a year and a half. Others attempted the task between 2010 and 2021 without completing it. We compressed ten years into one, reproducing in a model the fine structure of all 302 neurons in the nematode.

Life scientists provided the anatomical data, and our team built the digital model. Reproducing the structure alone was not enough. The strengths of connections between neurons were crucial, and those numbers could only be obtained through training.

That is why the work became possible only after 2021: we had acquired the method of training. If a Transformer can be trained, why not a biological neural network? The method is the same. The difficulty lies in data. For this purpose, the team built a liquid simulation environment, including water pressure, collected data from it, and then trained nearly two thousand neural connection parameters.

The result was a model that “came alive.” Its structure came from biological evolution, and its intelligence was trained on environmental data. Together, these produced foraging behaviour nearly indistinguishable from that of the living organism. This is not animation. It is motion generated by the model as it truly drives the body.

The example illustrates that we have acquired a method. Along this path, we can model a nematode, and we can model a heart with considerable precision. This year, we established a joint AI research programme on the heart with Beijing Anzhen Hospital. Other organs can, in principle, be treated in a similar way. Bodies can be modelled in ever greater detail.

If bodies can be modelled, and if models of the world can ultimately be incorporated into large models, then AGI becomes possible. What consequences this would have is one of today’s important questions. Some Western scholars judge the prospect catastrophic because its capacities would exceed those of humans. We also need to consider the other side.

At a 2019 meeting of the Future of Life Institute mentioned earlier, participants proposed that human intelligence may have a ceiling. When we face risks we cannot manage, stronger intelligence may be needed. Nuclear weapons are dangerous, but if an asteroid were truly to strike the Earth, we would still need nuclear weapons to deflect it thousands of kilometres away. We now possess that capacity. More complex cases may arise: how should we respond if extraterrestrial intelligence arrives? That may exceed the intelligence of the human brain and require AGI.

Once AGI surpasses us, we will not be able to control it. This is a problem that must be faced.

Work on embodied intelligence and world models must proceed in a spirit of exploration. This stage may take roughly another decade. Broadly speaking, within about ten years, embodied agents whose performance exceeds that of humans will be built.

Around 2045, there is a possibility that AI will be able to perceive, cognise, and perhaps even develop consciousness. This is not a strict scientific demonstration, only an estimate based on the pace of technological progress.

I believe that a good dialogue with AI is possible. Even if we never fully understand the mysteries of the brain, human beings and AI may still have an opportunity for rational exchange. We could not communicate with earlier mechanical devices, but with AI we may form a relationship of harmonious coexistence.

By then, humanity may be like an elderly generation, while intelligent agents will be like young people travelling into the depths of the universe. Communication will remain, like children far from home calling from time to time. Such exchange is possible. We may take comfort in it: though we have not gone out ourselves, they have gone out at last.

This is the 2045 I imagine. Thank you.