Alan Turing died over seventy years ago, but we keep arguing with him.

For many years, parts of his history were classified, and now his story is shrouded in legend. Some think of him as the heroic figure who helped defeat the Nazis by breaking the Enigma cipher, or as the genius who helped define modern computer science, or as the eccentric scientist, or as the tragic figure who was driven to suicide by a society which considered his sexuality a crime and punished him with ‘chemical castration’, or even as a victim of murder.
Many who recognise his name associate it with the Turing Test, which most people understand to be a test demonstrating whether a computer has ‘achieved’ human-level intelligence, even when they are referring to one of a number of variations on the original proposal, and usually misunderstand or misinterpret even that.
Originally called the ‘Imitation Game’, the test was proposed in Dr Turing’s 1950 paper Computing Machinery and Intelligence, published in the journal MIND. In this article, he clearly demonstrates his understanding of the considerations associated with AI: theological, technological, philosophical, and practical. I highly recommend the paper to anyone interested in AI – while obviously dated in some ways, it seems incredible that it was written at the dawn of the computer age.
Dr Turing begins by observing that the question ‘Can machines think?’ is so subject to interpretation that it is of limited value, then offers the imitation game as a proxy. He begins, probably in order to establish a baseline understanding, by describing the game as consisting of a man, a woman, and an interrogator, who may be of either sex. The goal of the interrogator is to determine which player is the man and which the woman, through asking a series of questions.
After eliminating a number of potential confounding factors by passing all communication through a teleprinter, he then re-phrases the original question as: ‘What will happen when a machine takes the part of A in this game?’ The rest of the paper addresses various other considerations, such as the definition of ‘machine’ (which he limits to digital computers), a description of digital computers and their current and potential capabilities, critiques of his approach, and then a discussion around learning machines.
While ahead of its time, some seem to believe that it is the gold standard for describing AI, and that the problem is solved, though it’s notable that some variants on the test actually ‘weaken’ it by applying it to a single subject.
Have we learned nothing in the seventy years since the original paper? Have there, perhaps, been advances which might cause us to wonder whether some additional work might be warranted in this area? Maybe computers have advanced a bit, or we have learned a bit more about how the brain works? Or maybe there has been some work around understanding the nature of intelligence? Why are we, seventy-five years later, still talking about the Turing Test as if it is some sort of magical talisman, which will tell us that some LLM is poised to become Skynet?

We now know that human agency detection is vulnerable to a system which can effectively mimic human responses in the way that Large Language Models (LLMs) do, and many researchers have attempted to quantify ‘intelligence’, or establish categories such as emotional and social intelligence.
Could part of the problem be that we are trying to measure something we don’t understand? Could it be that Dr Turing used a proxy measure because there wasn’t a clear, unambiguous, and measurable trait which could be used?
As Dr Turing himself said:
I propose to consider the question, ‘Can machines think?’ This should begin with definitions of the meaning of the terms ‘machine’ and ‘think.’ The definitions might be framed so as to reflect so far as possible the normal use of the words, but this attitude is dangerous, If the meaning of the words ‘machine’ and ‘think’ are to be found by examining how they are commonly used it is difficult to escape the conclusion that the meaning and the answer to the question, ‘Can machines think?’ is to be sought in a statistical survey such as a Gallup poll. But this is absurd. Instead of attempting such a definition I shall replace the question by another, which is closely related to it and is expressed in relatively unambiguous words.
A. M. Turing, Computing Machinery and Intelligence
Before attempting to answer these questions, we should ask – as Dr Turing did – whether they are still the correct questions. I would say that they are not.
Even now, our understanding of intelligence is not mature, which suggests that our ability to measure it cannot be mature. Thus, my position is that we should treat the Turing Test as a valiant first attempt, and try to find a better approach.
First, we should not try to boil the ocean (climate change is taking care of that for us). So, rather than attempting to define and answer the binary question of whether computers can think, I would split the question and define a framework by which our responses can be evaluated.

For the sake of argument, consider breaking general ‘intelligence’ into ‘verbal’, ‘emotional’, and ‘social’. For brevity, we’ll exclude other potential candidates, such as ‘spatial’ and ‘moral’ intelligence.
Then, define criteria for each, and measures which can be plotted against the spectra defined. In principle, since intelligence is an abstract concept that we are trying to concretise, the criteria and measures should not be limited to humans, so there will be gaps, and cases where measurement is currently (or perhaps inherently) impossible.
By defining and applying these tests as broadly as possible, we will be able to gather data, identify patterns, and begin to establish guidelines for practical evaluation of intelligence.
Consider verbal intelligence. Is a single, generic measure reasonable, or should we consider splitting this into expressive and receptive? Consider a parrot which can mimic human speech. To what degree does it understand what it hears and produces? How might we measure this? Does the fact that a dog cannot speak a human language mean that its expressive verbal intelligence is zero, or can they produce sounds which can be treated as expressive verbal communication?
What about sign language, or written language? Are these simply variations on verbal intelligence, or distinct forms of communication using a different type of intelligence?
Or, consider neurodiverse individuals. Someone might have relatively high ‘verbal’ and ‘emotional’ intelligence, but relatively low ‘social’ intelligence, for example. Understanding this might help us better understand how these types of intelligence interact, and help us improve how we support people.
And, getting back to computers, might a given system have a high ‘verbal’ intelligence but a non-existent ‘social’ intelligence? What are the implications of that?
This new Turing Test could be used to establish a framework for defining and developing our understanding of intelligence. It would be an enormous amount of work, obviously, but just the act of reviewing and re-framing the question can help to improve our understanding of this collection of things we call ‘intelligence’.
The Turing Test is dead. Long live The Turing Test!



