We still haven’t passed the Turing test
Is AGI already here, or are we still a while off? Demis Hassabis is adamant that we are still 5 to 10 years away. Then today, Dario Amodei reiterated that it should be here next year.
My measure for AGI progress is the Turing test, and I think that it shows that we are still a while off. Others would disagree, or would argue that the Turing test is no longer relevant to what is noteworthy about intelligence.
Here I’ll give my view on why the Turing test hasn’t been passed, and argue for its continued relevancy.
Does ChatGPT pass the Turing test?
In 2024, Peter Thiel suggested that the Turing test had mostly been passed:
The Turing Test means that a computer can fool a human into thinking that it is a human being. It has always been a somewhat fuzzy test: does it fool an expert or a non-expert? Does it fool you all of the time or only some of the time? How convincing is it? But to a first approximation, we were not even close to passing the Turing Test in 2021. Then ChatGPT basically passed the Turing Test—at least for an average person with an IQ around 100. It did pass the Turing Test. And that was the holy grail. That was the holy grail of AI research for the previous sixty years. — Peter Thiel (28:00)
I disagree, or interpret the advent of ChatGPT in a different way.
The percentage of time that AI fools a human jumped up dramatically with ChatGPT, from essentially nothing to a surprising amount of time. One way to interpret this is to say that the “Turing test has finally been passed (in some cases)”; the other way is to say “we are finally making progress towards passing the Turing test”. I take the latter view.
Simply, even an “average person” can tell they are talking to an AI within a minute of a phone conversation. It takes longer over chat, but then this is only initially. After repeated engagements, people become attuned to signs of AI, and then the timeframe before detection begins to drop back down.
Thiel’s argument then is that what was important about the Turing test had been passed — and I think it’s for this reason that plenty of researchers no longer consider the Turing test to be relevant. Going from “never convincing” to “convincing a lot of the time” meant that the important part of the test seemed to be cleared. But I think it’s the other way around. There are now table-stakes to passing the test, and its utility is clear: it helps us distinguish those things that are more or less integral to being generally intelligent.
In the case of ChatGPT, I think the Turing test proves that the ability to manipulate language is not the ‘holy grail’ beyond which you have replicated human intelligence:
This is another way of getting at what’s so crazy about ChatGPT passing the Turing Test. If we had sat here two years ago and you’d asked me what the distinctive feature of a human being is—what makes someone human in a way that differs from everything else—my go-to answer would have been language. You’re three years old, you’re eighty years old — just about all humans can speak languages, and just about all non-humans cannot. It’s this binary thing. — Peter Thiel (1:21:00)
Clearly, being able to manipulate language doesn’t get you everything you need to be generally intelligent in the way that humans are. Plenty of researchers prior to ChatGPT thought that it did, and plenty of researchers disagreed: it turns out that the latter were correct, with the Turing test helping to settle the score.
What it takes to pass the Turing test
The iteration of the test that I currently use is: can an AI call me up on my mobile, talk to me for an hour, and at the end of the conversation leave me unable to figure out whether I was talking to an AI?
There are potentially two objections here. The first is “you can tell it’s AI because it’s not trying to hide it”. The second is that voice chat isn’t a fair domain for the test — that Turing’s original test was limited to text.
On deception
It is a common misconception that the Turing test revolves around deception. It’s not supposed to.
Turing’s thesis was that being able to hold a human-like conversation requires human-like intelligence — the two were equivalent. To imitate a human in conversation requires replicating what’s notable about human intelligence.
That ChatGPT can imitate human conversation some of the time suggests we are on-the-money with deep learning techniques. It is likely that we are replicating important features of the human brain to be able to get AI to be as convincing as it currently is. But that we cannot do the phone call yet, or that AI fails over longer text chats, shows that we are still missing some important features.
A related argument is that AI only fails because of irrelevant giveaways inherent to the form of, rather than the content produced by, the AI system. For example, that a LLM would struggle to convey a fabricated backstory of a human life, because it does not have these memories stored in the exact same way the human brain does, and so you can sort of tell that it’s not being ‘recalled’ in the same way the human brain does (and that this type of memory recall is beside the point to AGI).
I think this is an irrelevant argument, and there are likely ways to structure the test so that it avoids these giveaways. For instance, the human that is being tested is also given a fabricated backstory. Or — and more on this later — just let the tester know that it is talking to an AI ahead of time.
The test is not supposed to be about deception, so it shouldn’t be relevant whether you know you are talking to an artificial system ahead of time. (Think of the movie Her).
On real-time chat
Simply, it’s easier to measure how long a conversation takes over voice than over text.
A similar objection to before is that this introduces irrelevant signs of artificial intelligence that is beside the point to the content of the conversation. The first sign might be “no tacit feel for the use of vocal cords”. The second sign might be “no tacit feel for human-specific latency”.
I think these are reasonable challenges for AI to overcome; I also imagine that AI companies will solve for both of these before they solve for the Turing test. But either way, you could just introduce a regularising voice filter for both human and AI speakers, and some randomised amount of latency. These interventions shouldn’t impact the test greatly over the course of an hour (i.e., humans adjust quickly to latency and distortion over the phone, which seems strongly related to what makes us generally useful in the world).
On emotion, and humour specifically
Today, and I believe this will hold up for while, the fastest giveaway that you are talking to AI is that it cannot be funny in a conversational setting. This is much more challenging in real-time, as opposed to text, which is why I think real-time is a better setting to test in.
What’s so particular about humour that AI can’t do it? I get the sense that some researchers wave it off as “possessing particular faculties of human emotion that are unrelated to manipulating the natural world”. I tend to disagree, although I’m not entirely clear on why, and will leave it for another day.
(While humour tends to be the easiest to pick up on, there are other emotions that I do not think are “specifically human and irrelevant to general intelligence”. Boredom and anxiety seem partly human-specific, but also partly universal to generally intelligent beings, because they are words used to discuss internal states that correspond to reasoning under uncertainty, limited resources and time-constraints in the real world).
What it is that is being tested for
Given that the Turing test is not about deception, I will revise my earlier framing of the test: can an AI call me up on my mobile, talk to me for an hour, and at the end of the conversation leave me thinking that I’ve spoken to a person?
The movie Her is illustrative: the main character knows he is talking to an AI, but it doesn’t matter. The AI is as much a person to him as any of the people around him. As viewers, we also feel this.
Personhood has its own challenges. Some people think Claude is a person. Some people think that statues, volcanos or trees are people. And then sometimes people don’t think certain other people are people (e.g. Aristotle and the barbarians, ironically because he thought speaking Greek was a sign of personhood).
While there are these edge cases, I believe the essence of the Turing test is clear enough. Can I talk to you and feel like I am talking to a friend or acquaintance? It seems trivial, but we are not there yet with AI.
Why the test has fallen out of favour
You’d imagine that as we get closer to passing the Turing test it would become more talked about, rather than less. But I think there’s a simple explanation for why this happened: as we finally got empirical examples to conduct the test on, we become more unclear on what it would take to definitively pass the test. This is the same as saying we became increasingly unclear as to what it is that makes a person a person.
Framed another way: as AI becomes more humanlike, what it is to be human becomes less clear. That’s what Peter Thiel was observing: a lot of people, himself included, equated language-use to humanity, which turned out to be incorrect. What then are the defining features of humanity? The answer isn’t on the tip of the tongue like it used to be in the 00-10s: “Oh, we’ll just know”.
I’ve argued however that really the opposite should be true. The Turing test is only now relevant, because it serves as a handy tool to keep us focused on a strong qualitative threshold of general intelligence, as opposed to a broad bench of quantitative measures (each of which is deemed to demonstrate AGI, before quickly being dismissed after being passed).
The Turing test’s strength is that it is qualitative: what it takes to pass the test changes over time, as people’s measure of personhood, much of which is tacit, changes over time. Bring someone from 2006 to the present-day, and they will immediately think they have AGI in their hands when they pick up ChatGPT. Clearly the signs that we were preemptively looking for were not, in fact, the signs (this also explaining how ideas of general intelligence have changed significantly throughout the course of human history itself).