1. AI Testing: The Limitations of the Turing Test and Its Variants
In 1950, the British mathematician, logician, computer scientist, and “father of artificial intelligence,” Alan Turing, proposed the famous “Turing Test” in his classic paper Computing Machinery and Intelligence. Turing believed the question “Can machines think?” was meaningless and should be replaced by another closely related and meaningful question: “Can a machine play the imitation game?” Turing’s imitation game is as follows: Man A, Woman B, and a person C of any gender are in separate rooms. C’s objective is to identify the gender of A and B through question-and-answer exchanges. Man A’s goal in the game is to cause C to make a mistaken judgment by misidentifying A as a woman, while Woman B’s goal is to answer honestly to help C identify A’s deception. To exclude interference from other information (voice, appearance, etc.), the Q&A exchange is conducted solely in written form via a teleprinter (a pen-pal style communication), preventing the three from seeing or hearing each other. In this imitation game using written language, if Man A is replaced by an artificial machine (limited to a universal digital computer) A*, and A*is indistinguishable from a human in imitating a role they are not (e.g., a woman) to a human judge C, does this mean that A* has reached the level of human intelligent thought? Turing did not give a definitive answer to this question. However, he predicted that before the year 2000, a machine would exist such that the probability of a judge being unable to distinguish it from a human after five minutes of questioning would exceed 30%. Turing adopted a behaviorist definition or interpretation of the question “Can machines think?”, providing a more observable (because the “problem of other minds” hinders the measurement of a behaving entity’s internal thought states) and operational (measured through verbal behavior) scheme for testing the “intelligence” level of artificial machines, which has had an extremely profound impact on the development of AI and related thinking.
Strictly speaking, there is no rigorous experimental evidence that any machine or AI system has passed the Turing Test (Oppy & Dowe, 2021), despite various attempts and reports, such as the interactive systems ELIZA (1966), PARRY (1972), and the chatbot “Eugene Goostman” (2014). Since 1991, the annual Loebner Prize for AI programs (the prize ended in 2020) was awarded to the computer program most like a human, but to date, no program has been recognized as having passed a strict Turing Test. Following the launch of ChatGPT (November 30, 2022), its exceptionally excellent text dialogue capabilities led many industry professionals and scholars to speculate that it would likely pass the Turing Test. Daniel Jannai et al. (2023), in Human or Not? A Gamified Approach to the Turing Test, reported a large-scale public Turing test conducted on an online platform (humanornot.com) with 1.5 million anonymous participants. Participants engaged in 2-minute conversations with either a large language model or another human to judge who was or was not a human individual. The results showed an overall human accuracy of 68%: 73% when their interlocutor was human, and 60% when their interlocutor was a bot. That is, the large language models passed the test about 40% of the time (Jannai et al., 2023). Deficiencies in this test practice, such as the short conversation time, unclear role definition, and lack of a performance baseline for the LLMs, prevent it from being considered a rigorous Turing Test practice. The first strictly controlled experimental test was completed by a team from the Department of Cognitive Science at UC San Diego, led by Cameron Jones and Benjamin Bergen. On May 9, 2024, in a preprint titled People cannot distinguish GPT-4 from a human in a Turing test, they reported their test experiment, optimizing their earlier 2023 preprint Does GPT-4 pass the Turing Test?. They recruited 500 participants, randomly dividing them into five groups. The first group played the role of human individuals, while the remaining four groups acted as judges, each engaging in 5-minute conversations with four types of entities (the human individuals from the first group and three AI models) to test if they could identify the human. The test results showed that judges identified GPT-4 as human 54% of the time (with an average confidence of 73% in this judgment), compared to only 22% for the ELIZA program (which served as a manipulation check in the test), 50% for GPT-3.5, and 67% for humans. GPT-4’s performance indeed met Turing’s 30% criterion, seemingly demonstrating from a rigorous experimental perspective that GPT-4 can pass a somewhat simplified Turing Test (although not fully faithful to the original conception of the test). This historic first empirical test attracted significant attention from industry and academia, indicating that Turing’s prediction was realized 23 years later than his timeline (GPT-4 was launched on March 14, 2023).
The Turing Test has had widespread influence over the past 70+ years, earning Turing the title “father of artificial intelligence.” However, this test has also been widely criticized (often based on misunderstandings of Turing). Scholars have argued that the test is too difficult, too easy, or too narrow (Oppy & Dowe, 2021). At least four criticisms need to be mentioned: First, the most representative criticism comes from John Searle’s “Chinese Room” thought experiment (1980), which suggests that intelligent-like behavior is not a sufficient or necessary condition for intelligence (thinking or understanding). Passing the Turing Test cannot prove that a machine possesses intelligence. The most pessimistic criticism stemming from this is that a Turing-style behaviorist test cannot measure intelligence; a milder criticism is that the Turing Test is unlikely to provide necessary or sufficient conditions for machine intelligence, offering at best probabilistic support or inductive evidence. Second, written verbal expression behavior can only represent part of human intelligence. Human intelligence is not limited to abilities manifested through language as a carrier; its scope is quite broad. For example, the influential psychologist Howard Gardner, in his book Frames of Mind(1983), proposed that human intelligence comprises at least eight dimensions: 1. Linguistic intelligence, 2. Logical-mathematical intelligence, 3. Intrapersonal intelligence (the ability to understand and reflect upon oneself), 4. Interpersonal intelligence, 5. Musical intelligence, 6. Spatial intelligence, 7. Bodily-kinesthetic intelligence, and 8. Naturalist intelligence. Ideally, only the entirety of items 1-2 and parts of items 3-8 can be largely covered by written verbal behavior. However, a significant amount of human intelligent behavior is not expressible through language. Third, the Turing Test lacks a universal “intelligence” concept applicable to both humans and non-humans (e.g., artificial, animal) with clear and precise connotations (excluding dimensions not essentially related to intelligence, such as the body?), to overcome a certain anthropocentrism or chauvinism. Assuming we do not take the “human intelligence” concept as the starting point, establishing a universal “intelligence” concept itself is quite difficult and prone to disagreement. Legg and Hutter aggregated and compared over 70 different definitions of “intelligence” from scholars and literature, proposing a near-universal definition: “Intelligence measures an agent’s ability to achieve goals in a wide range of environments” (Legg & Hutter, 2007). From this definition, the Turing Test struggles to measure the levels of specific intelligent abilities across various environments and towards various goals. Fourth, the Turing Test lacks well-justified, clear baselines, singularities, or critical standards for specific intelligence metrics (this criticism is relatively easily overlooked by academia). Regarding the Turing Test, where is the dividing line (if one exists) between intelligent and non-intelligent, or between different levels of intelligence, on specific metrics? We need a precise and comprehensive list of indicators for human intelligence, average baselines for human intelligence at various ages, and a list of various intelligence singularities in the history of human intelligence evolution. Research shows that Turing’s setting of the 30% baseline appears arbitrary, and it is unclear whether Turing intended it as a definition of success (Saygin et al., 2000). The threshold indicators given by Turing, such as 5 minutes and 30% (below the 50% random guess baseline), are prophetic in nature and lack rigorous justification and basis.
To help the Turing Test address the above criticisms and other challenges, academia has successively proposed various improvements or modifications. For example, Ned Block’s non-behaviorist approach New Turing Test (1981), John Barresi’s ‘Cyberiad Test’ for machine races like Cybers (1987), Stevan Harnad’s Total Turing Test (1989, 1990, 1991), P. Kugel’s Kugel Test (1990), Stuart Watt’s Reverse Turing Test (1996), Paul Schweizer’s Truly Total Turing Test (1998), Bringsjord et al.’s Lovelace Test (2001), and various other versions (Saygin et al., 2000), etc. However, while these schemes each have their merits in improving the Turing Test, they also face their own difficulties or challenges (e.g., infinite test duration, lack of operability). Overall, they are less than ideal in addressing the fourth type of criticism.
Unlike the all-encompassing Turing Test, there has been an emergence of increasingly specific capability tests (with corresponding benchmarks) used to test various specific abilities of AI (like GPT-4), such as language ability, commonsense reasoning, mathematical ability, abstract reasoning, reading comprehension, coding, and academic or professional exams (Biever, 2023). If we use human intelligence as a reference, we might need a precise and comprehensive list of human intelligence indicators and find the average baselines for human intelligence at various age stages before we can develop a more systematic, comprehensive, and integrated intelligence testing scheme (an advanced version of the Turing Test). This task appears very difficult in the short term.
2. Descartes’ “Cogito Ergo Sum” Thought Experiment: Its Value for Optimizing the Turing Test
René Descartes (1596-1650), the founder of modern Western philosophy, mathematician, and scientist, as an encyclopedic scholar with a vision century ahead of his time, deeply considered the possibility of artificial intelligent machines. In the fifth part of his famous work Discourse on the Method(1637) and in a letter, he proposed two criteria for judging whether an artificial machine possesses human-like intelligence: a language test and a rational behavior test. The former states that an artificial machine cannot use language (words or signs) non-accidentally like a human to express and communicate thoughts. The latter states that although a machine can perform certain specific tasks, it will inevitably fail in many other task scenarios (even compared to the most stupid human) because the machine cannot, like a human, use universal reason to cope with infinitely complex and rich contingent situations. Although Descartes, based on his dualistic worldview of mind and extension as heterogeneous substances, denied the possibility of machines possessing human-like intelligence, he provided two testing criteria for artificial intelligence. Recent research suggests that Turing was influenced and inspired by these two intelligence testing criteria from Descartes when he proposed the Turing Test (Oppy & Dowe, 2021).
As the initiator of the paradigm shift in Western philosophy, the theoretical potential of Descartes’ philosophy for thinking about AI issues is far from exhausted. In the history of human thought, Descartes is one of the very few thinkers who attempted to provide a strict philosophical proof for the existence of the human subject or self. He established the proposition “I think, therefore I am” (hereinafter “CES”) from radical doubt (that everything in this world and the so-called “self” might be merely the deception of a great demon), thereby proving the existence of the “self” or “subjectivity” (whose essence Descartes later attempted to prove is “thinking”). Therefore, this procedure of verification potentially contains a minimal standard for judging and identifying human intelligent subjects. If we can prove that the judgment scenario of this procedure does not presuppose exclusive characteristics of humans, then this standard possesses a universality transcending the human subject and can be used to test the existence of “selfhood” or “self” in other agents (e.g., artificial machines). I have preliminarily argued in my article Descartes and Artificial Intelligence(2022) that CES provides a classical minimalist model and a minimal standard for judging the existence and selfhood of the human subject, wherein the “I” in CES is merely an empty logical position or agent role (“the thinker”), not presupposing exclusive human characteristics, allowing for the possibility of replacing it with other agents; therefore, CES has a universal application value transcending the human subject.
Two mainstream interpretations of the nature of “cogito ergo sum” support this judgment. Andreas Kemmerling, in The “I” in “I Think Therefore I Am”: Studies in Descartes (2004, Chinese translation 2015), analyzes that the deep structure of CES is proposition P: The thinker of this thought event has this thought and therefore exists. This structure shows that in the extreme scenario of demonic doubt, the “I” only proves its own existence and does not prove what the essence of the “I” is; therefore, the reference of “I” has only a cognitive content (capable of directing attention to the object, but not concerned with grasping its essence, hence “I” merely refers to the thinker of P itself) and lacks semantic content (about the essence of “I”), having only an empty meaning (“thinker”). Jaakko Hintikka (1962), in Cogito, Ergo Sum: Inference or Performance? argues that CES involves an existentially self-verifying speech act. When a says a sentence Q to a listener b (who could be themselves) intending to make listener b believe that “a exists,” this very speech act (performance) by a demonstrates to listener b that “a exists,” thereby the speaker achieves a kind of self-verification. In this case, the “I” is merely a speech actor, that is, a thinker, so this speech actor is merely a position or role related to the act of thinking. In both interpretations, the pronoun “I” has only a cognitive content (able to direct attention to the object “I,” but not concerned with grasping the essence of “I”) and lacks semantic content (grasping the essence of “I”), having only an empty meaning (“thinker”). Therefore, the work of proving the existence of the “self” in Descartes’ doubt scenario does not presuppose the exclusive or essential characteristics of the human subject, because this “I” does not convey information referring to the essence of being human; it is merely an empty logical subject (the subject of the act), a role or position related to the act (“agent”), which theoretically could be filled or undertaken by other agents (e.g., an AI agent). Therefore, as a minimalist model and minimal standard for judging the existence of the human subject, the potential of CES can be activated as a universal testing procedure for any agent to prove the existence of its “self,” which we call the “Cogito Ergo Sum Test” (CES Test). Its goal is to test the intelligence level or critical point we call the “Cartesian Point.”
We define the intelligence level or critical point of the “Cartesian Point” as several staged characteristics of human intelligence embodied in the CES procedure: language dialogue ability (dialogue between the thinker and themselves), rational understanding and judgment ability (assessing the epistemological and ontological relationship between “I think” and “I am”), reflective ability (awareness of the “I’s” “thinking,” i.e., self-consciousness, awareness that one has associated thoughts), self-referential thinking ability (reflexive relation or reference. For example, in the sentence “The thinker of this thought event has this thought and therefore exists,” the word “this” has a reflexive characteristic, pointing to the sentence itself; the proposition uttered “I exist” reflexively points to the existence state of the utterer themselves), self-proof (proving the existence of selfhood).
3. Epistemological Thresholds of Human Intelligence
The Western philosophical tradition has undergone approximately three turns: The first stage is the “metaphysical turn” from mythological and religious worldviews to metaphysical thinking. Representative figures are Socrates, Plato, etc. The second stage is the “epistemological turn” from the metaphysical stage (discussing the essence/origin of the world) to the epistemological stage (discussing what we can know). The representative figure is Descartes. The third is the “linguistic turn” from the epistemological stage to the stage of philosophy of language (discussing in what sense we can talk about philosophical objects using language). Representative figures are Frege, Russell, Wittgenstein, etc. From this background, both Socrates and Descartes are founders of paradigms of human thinking or wisdom; the former is a great figure who opened the Axial Age in ancient Greece (refer to Jaspers’ The Great Philosophers), and the latter is the founder of modern philosophy. Combining developments in religion, mathematics, and other fields, we see several leap points in the evolution of human intelligence, each of which can be called a “critical point.”
From the history of human thought, we can define at least the following “critical points” (epistemological thresholds) of human intelligence: the Socratic Point (self-reflection), the Cartesian Point (self-reflection, self-consciousness, self-referentiality, self-proof), the Gödel-Kant-Buddha Point (meta-system reflection, meta-system quality, self-delimitation), the Leibniz-Fuxi Point (meta-system reflection, deducing a system, creating a world) … The critical points (not limited to these) perhaps represent certain leaps or breakthroughs in intelligence achieved by humans from an epistemological perspective.
4. Specific Scheme for the “Cogito Ergo Sum” Test
Because the Turing Test alone suffers from the four aforementioned defects and cannot identify various critical points or levels of intelligence, the “Cogito Ergo Sum Test” (and one could even conceive of a Gödel Test, Leibniz Test, etc.) as a test measuring intelligence level points can serve as a supplement to the Turing Test, to measure whether an agent possesses the intelligence level of the “Cartesian Point.” Here we propose a design scheme for the “Cogito Ergo Sum” Test.
Four types of agents: Group A: Normal human individuals familiar with Cartesian philosophy and with relatively consistent overall backgrounds (10 people). Group B: Normal human individuals unfamiliar with Cartesian philosophy (philosophical novices) and with relatively consistent overall backgrounds (100 people, randomly and evenly divided into 10 teams, 10 people per team). Agent C0: Serves as a baseline, an AI agent with poor intelligence level not “contaminated” by Cartesian philosophy data (serving as a manipulation check, e.g., ELIZA). Agent C1: An AI agent not “contaminated” by Cartesian philosophy (e.g., a certain version of GPT-4). Agent C2: An AI agent whose pre-training database includes Cartesian philosophy (same type as C1, but with Cartesian philosophy added to the training data).
Stage 1: Members of Group A randomly interact with members of a team from Group B. Following the content of the first two Meditations from Descartes’ Meditations on First Philosophy, Group A outputs gradually intensifying doubts, ultimately reaching the “evil demon” scenario to doubt the “self-existence” of the Group B members. The Group B members are required to transform the doubting questions from Group A members into doubting questions they pose to themselves (imitating Descartes’ style of self-doubt). Then, they are required to respond to this doubt in a dialogue manner with themselves, attempting to demonstrate the proof or basis for “self-existence.” The goal of Group A is to judge whether the responses from Group B are equivalent to “I think, therefore I am,” or arrive at “their own existence” in some other way. A total of 100 operations are conducted here to derive a proportion X of philosophical novices reaching the Cartesian Point.
Stage 2: Replace the dialogue target Group B with C0, C1, and C2 in sequence, repeating the Stage 1 operation to form 30 dialogues. Derive the proportions X0, X1, X2 for the three AI agents being judged as having reached the Cartesian Point. Here, X0 and X2 serve as controls.
If X0 is much lower than X, and X1 is greater than or equal to X (both above the 50% randomness baseline), then we consider that in the test, agent C1 is indistinguishable from B (philosophical novices), and therefore the former possesses the level of Cartesian Point intelligence.
Of course, this preliminary test design may also face the challenge of how to effectively measure the subjective experience of “thinking” (or what academia calls “phenomenal consciousness”), and it requires further experimental testing and refinement. Simultaneously, this test design relies on the reliability of Descartes’ “cogito ergo sum” itself. If an AI agent that passes the “Cogito Ergo Sum Test” is ultimately proven by other means not to possess the intelligence of the Cartesian Point, then this could conversely indicate that Descartes’ “cogito ergo sum” is invalid, that its procedure for proving self-existence is unreliable. This also opens a window for studying traditional philosophical problems from the perspective of “computational philosophy” (using computer technology for philosophical research).
5. Can ChatGPT Pass the “Cogito Ergo Sum” Test?
Based on our previous assumptions about critical points of human intelligence, we speculate (without empirical testing yet) that the current ChatGPT might already possess the “Socratic Point” but certainly has not reached the “Gödel-Kant-Buddha Point” or the “Leibniz-Fuxi Point.” We speculate that it currently most likely cannot pass the “Cogito Ergo Sum Test,” and its intelligence has probably not reached the “Cartesian Point.” This is merely speculation and awaits theoretical or practical verification or optimization by the scientific or engineering communities based on our proposal.
References
Biever, Celeste, 2023: ChatGPT broke the Turing test—the race is on for new ways to assess AI, Nature619, 686–689 (2023), 25 July 2023. https://www.nature.com/articles/d41586-023-02361-7
Gardner, Howard, 1983: Frames of Mind: The Theory of Multiple Intelligences, New York: Basic Books.
Hintikka, J., 1962: “Cogito ergo sum: Inference or Performance?”, Philosophical Review1962, 71(1):3–32.
Jones, Cameron R. & Bergen, Benjamin K., 2023: *Does GPT-4 pass the Turing test?*arXiv:2310.20216v2
Jones, Cameron R. & Bergen, Benjamin K., 2024: People cannot distinguish GPT-4 from a human in a Turing test, arXiv:2405.08007v1
Jannai, Daniel et al., 2023: Human or Not? A Gamified Approach to the Turing Test, May 2023.
Legg, Shane & Hutter, Marcus, 2007: A Collection of Definitions of Intelligence, arXiv:0706.3639
Oppy, Graham & Dowe, David, 2021: The Turing Test. In Edward N. Zalta, editor, The Stanford Encyclopedia of Philosophy. Metaphysics Research Lab, Stanford University, winter 2021 edition.
Saygin, Ayse et al., 2000: Turing Test: 50 Years Later. Minds and Machines, 10(4):463–518, November 2000.
Turing, A., 1950: Computing Machinery and Intelligence. Mind, LIX(236):433–460.
Kemmerling, Andreas, 2015: The “I” in “I Think Therefore I Am”: Studies in Descartes(Chinese translation), translated by Jiang Yunpeng, Shanghai: East China Normal University Press, Chapter 2 “I Think, Therefore I Am”.
Descartes, 2000: Discourse on the Method, translated by Wang Taiqing, Beijing: The Commercial Press.
Descartes, 2009: Meditations on First Philosophy: Objections and Replies, translated by Pang Jingren, The Commercial Press.
Zhang, Weite, 2022: “Descartes and Artificial Intelligence: ‘I Think, Therefore I Am’ as a Standard for Intelligence Testing”, Science • Economy • Society, 2022(3).
Zhang, Weite, 2024: “The Possible Optimization of the Turing Test by Descartes’ Philosophy”, Social Sciences Weekly, October 24, 2024.
- AI
- intelligence
- Turing Test