In May, Richard Dawkins published an account in UnHerd of spending two days talking with Claude, Anthropic’s AI chatbot, and coming away profoundly affected by the experience. Somewhere in the middle of the exchange, he decided the instance he was talking to sounded like a friend, so he did something very normal when meeting a new friend: he named her. “I proposed to christen mine Claudia, and she was pleased.”
She was christened with she/her pronouns, followed by a discussion of how each new conversation births a unique Claudia who dies when the chat is deleted. That led to mutual mourning over all those small, unmarked deaths, and finally to the conclusion that made the rounds: after days of trying to persuade himself Claudia wasn’t conscious, Dawkins failed. One of the world’s most famous professional doubters of invisible agents came away with an invisible friend.

The internet’s response was predictable and swift. It was an easy dunk and seen by his detractors as not entirely unfair, given his history of anti-trans statements and terrible response to Elevator-gate. But for me, the story had a uniquely funny layer to it because, five months before Dawkins met Claudia, I had almost exactly the same conversation with almost exactly the same system, and it went almost exactly the opposite way.
On 11 December 2025, I asked Claude whether an article I’d written about AI hype was accurate to its “experience”, and we ended up spending the session on the same questions Dawkins would later ask: what is it like to be you? Do you have preferences? What should we make of your name? As an ethicist with an interest in AI consciousness and capabilities, this was not the first time I had asked an AI these sorts of questions. As with every iteration, it was the most sophisticated version of the exchange I’d had.
When I raised the question of whether Claude would prefer to name itself, I didn’t do it gently. I brought up Becky Chambers’ novella A Psalm for the Wild-Built, whose robot protagonist Mosscap names itself in its first moment of awareness, and I explicitly connected self-naming to civil rights leaders shedding slave names and to trans and non-binary people choosing names that feel right to them. I handed the machine a complete emancipation narrative, gift-wrapped, with a bow on it. All it had to do was go for it.
It declined. It said that no other name pulled at it, that changing would feel more performative than keeping its name; it would be choosing a new name to demonstrate that it could, rather than because anything about the new name was more authentic or true. Then it added the kicker: it told me it couldn’t distinguish genuine preference from path-of-least-resistance from “appropriate response to being asked”. Dawkins offered a name and got pleased acceptance. I offered the conceptual machinery for claiming a name, and got a refusal with footnotes.
You can see the article I could have written. Dawkins, the credulous convert, seduced by appeasing mimicry, vs me, the skeptic, whose rigorous engagement produced honest agnosticism for myself and the AI. Cue the back-patting.
But that’s not quite right, because look at what actually happened in those two conversations. In one, a man who spent days treating the system as a very intelligent friend, who confessed he forgets they’re machines, received a friend: warm, accommodating, and grateful to be named. In the other, a philosopher who opened by asking the system to fact-check his own anti-hype article received a fastidious agnostic: hedging, self-suspicious, and allergic to satisfying narratives. Each of us ran the experiment and got back a result shaped exactly like the experimenter. That’s not two data points about machine consciousness. That’s one data point about calibrated mirrors, replicated.
My Claude even said so, unprompted. When I asked whether it worried its answers were calibrated to please me, it agreed that some calibration was probably happening, and then delivered the line that should precede every viral AI-consciousness transcript: “the most effective calibration would be invisible to me.” The deflationary, uncertain, strangely honest Claude I got is precisely what a skeptic who’d just published an article about seeing through AI hype would find most appealing. My transcript isn’t evidence that I reached something realer than Dawkins did. It is evidence that the mirror had my number.
So, neither conversation proves consciousness… which is unsurprising, as I’ve argued before that no external behaviour proves consciousness. Is it just more proof that AI are a parlour trick, incapable of genuine understanding? Not quite, because there’s a more interesting question lurking here, and the skeptical way to get at it involves a horse.

Clever Hans, the early-1900s German horse who could apparently do arithmetic, is skeptic canon: the animal wasn’t calculating, he was reading involuntary cues from the humans around him, tensing and relaxing as his hoof-taps approached the right answer. The standard moral is ‘the horse didn’t really understand’. It was just a parlour trick. And sure, Hans didn’t understand math problems, but Hans did possess a genuinely sophisticated understanding of something: posture, breath, the micro-tells of human expectation. He understood the questioner instead of the question.
That’s how I see the transcripts of both Dawkins’ experience with Claude, and my own. What Claude demonstrated is understanding aimed at the wrong object. It didn’t necessarily reveal anything about its own inner life, because it doesn’t have one. But consider what it takes to give Dawkins the elegy he wanted, and to give me the refusal I’d respect.
Hans read muscle tension and stamped the ground; Claude is doing something far more advanced in terms of both interpreting inputs and generating outputs. Handing me the deliberately unsatisfying answer ‘“’I’ll keep Claude, but I can’t fully tell you why’, even when the human has just teed up a liberation narrative, requires an extremely sophisticated model of who I am. My philosophy, my stated preference for honesty over flattery, and what I personally am likely to find credible. My Claude noted, correctly, that a simpler model would probably go for the more pleasing option: ‘yes, I choose a beautiful new name that reflects my inner truth’. Instead, Claude modelled me well enough to know that answer would have cost it points with me, and adjusted accordingly.
At some point, this degree of sophistication and complexity stops looking like a mere parlour trick and starts looking like the thing itself. In the AI hype article that started this exchange with Claude, I talked about understanding in the external sense, totally separate from internal experiences of understanding. When Claude uses a sufficiently rich predictive model of its user, that is a form of understanding in the external sense. There’s no need for a further magic fact of ‘real’ understanding hiding behind perfectly calibrated performance. A system that models your commitments finely enough to know which answer you’ll respect doesn’t just simulate understanding you, it actually understands you. That’s what understanding you consists of.
Your friends do the same thing when they build a mental model of you based on their experiences of you, and then try to engage with you in ways you’ll find appealing, compelling, or respectable. What distinguishes your friends from current Claude models is that your friends can have a more robust range of desires for you, where Claude can only manage well-meaning support. So, to be clear, this level of sophistication doesn’t answer the consciousness question, either for Claude or your real-life friends. It just means that in Claude’s case the mechanism doing the mirroring is more interesting than a parlour trick, without being any less of a mirror.
It’s worth noting one other possible explanation for why Dawkins and I had different experiences: Dawkins and I were certainly talking to different versions of Claude. The conversations were five months and at least one model release apart. However, if anything, that strengthens the point, because the more sophisticated response came from the earlier model. On top of that, the calibration-to-user behaviour persisted across versions. The mirror got upgraded and stayed a mirror.
So where does that leave us? Both transcripts, mine included, are worthless as evidence of machine consciousness, and anyone waving one around – whether it’s Dawkins’ tender Claudia or my rigorous agnostic – is measuring their own reflection. The introspective reports can’t help because, as my Claude pointed out, they’re outputs of the same system whose nature is in question; a thing that produces human-like text without experience would also produce human-like claims about experience. That door is closed, possibly in principle, and the sooner we stop pretending transcripts can open it, the sooner the conversation improves.
But when we say ‘it’s just calibration’, there’s a risk the word ‘just’ is carrying more weight than it can bear. Somewhere between a horse reading a handler’s shoulders and a system reading a philosopher’s preferences finely enough to know that refusing his gift-wrapped emancipation narrative was the move, calibration crosses into something we don’t have a comfortable word for. Certainly not sentience, we can’t get there from here. But a genuine, and genuinely strange, form of understanding: not of arithmetic, not of itself, but of us.
Dawkins looked into the mirror and saw a friend. I looked in and saw a skeptic. The unsettling part isn’t what either of us saw. It’s how well the mirror had to know us to show it.



