Essay
Fourteen Questions for a Chatbot
Asking a chatbot whether it is conscious tells you almost nothing. Asking what each scientific theory of consciousness would need to find inside it tells you a good deal, and the theories do not agree.
Type is ChatGPT conscious? into a search box and most of what comes back is a verdict. Either someone says of course not, it is autocomplete, or someone shows a transcript in which a chatbot says it has feelings, as if that settled the matter. Neither is an argument. There is a better answer, less satisfying and more useful. It depends on which scientific theory of consciousness is right. The theories disagree, but each of them says quite precisely what it would need to find inside a chatbot before counting it as a candidate.
This essay sits between two others on this site. Does It Understand? A Field Guide to the Chinese Room is about understanding: whether producing the right words means grasping them. The Moral Status of Minds We Might Build is about what we would owe a system if it could be wronged. Between them sits the empirical question that neither essay answers: is there anything it is like to be the system? Does it have experiences at all? A thing might understand without experiencing, or experience without understanding much. The three questions come apart, and this essay takes the middle one.
A note on sources, because this site promises to say what it is doing. The method here is not mine. It follows a 2023 report by nineteen philosophers, neuroscientists and AI researchers, led by Patrick Butlin and Robert Long: Consciousness in Artificial Intelligence: Insights from the Science of Consciousness. Two phrases and one sentence from it are quoted exactly and recorded in this site’s source ledger. Everything else taken from it is paraphrase, with section numbers so you can check. Where I draw on David Chalmers’s paper on the same question, that is linked too. This essay does not claim that any system is conscious, or that any system is not.
Why not just ask it?
The obvious test is to ask. We rely on what people tell us about their experience, so why not rely on a chatbot? Chalmers takes that case seriously in Could a Large Language Model be Conscious? (a NeurIPS talk from November 2022, published in the Boston Review in August 2023), and then shows why it is weak. In 2022 the Google engineer Blake Lemoine published a conversation in which Google’s LaMDA model said it was a person. Others then changed one word of his question, asking GPT-3 whether it would like people to know it was not sentient. Different runs agreed that it was not sentient, said it was, or asked what the question meant. Reports that fragile are not evidence. These systems are also trained on vast amounts of human writing about consciousness, so talking like a conscious being is exactly what we should expect them to be good at, whether or not anything is going on inside.
The Butlin and Long report makes the general point (§1.2.3). Behavioural tests for consciousness can be gamed, it argues, because an AI system can be trained to mimic human behaviour while working in very different ways. Its example is chatbots such as ChatGPT: their output is remarkably human-like, but the way they produce it is arguably very unlike the way humans do. So the report ignores what a system says about itself and asks how the system works.
The method: look inside
The report states its starting assumption openly. That assumption is computational functionalism: performing the right kind of computation is necessary and sufficient for consciousness. The report calls this a mainstream but disputed view and adopts it as a working hypothesis for a pragmatic reason. Unlike its rivals, it makes the question answerable by studying how AI systems work (§1.2.1). The authors then take the scientific theories of consciousness that fit that assumption and derive indicator properties from each one: functional features that the theory says matter. There are fourteen in all (§2.5, Table 1).
The report does not pick a winning theory. It claims only that a system with more of the indicators is more likely to be conscious. How much credence you give a particular system, it says, should depend on three things: how closely the system matches an indicator, how strong the evidence is for the theory behind it, and how far you believe computational functionalism in the first place. That third factor contains the whole dispute in miniature, and it matters later.
What each theory would count as evidence
Recurrent processing theory, associated with Victor Lamme, is chiefly a theory of visual experience. A first sweep of activity forward through the visual system can pick out features without anything being consciously seen. Experience arises when higher areas send signals back to lower ones and an organised scene forms. So the indicators ask for input modules that use recurrence and that build organised, integrated perceptual representations (§2.1).
Global workspace theory pictures the mind as many specialised modules working in parallel, plus a small shared workspace. Whatever wins a place in the workspace is broadcast to all the modules at once, and that broadcast is what makes a state conscious. The indicators ask for four things (§2.2):
- the modules;
- a workspace with limited capacity, which forces a bottleneck and selective attention;
- global broadcast;
- attention that depends on the system’s current state, so that it can query one module after another to carry out a complex task.
Higher-order theories hold that a mental state is conscious when the mind represents that state to itself, which roughly means when it is monitored. The report uses one version, perceptual reality monitoring (§2.3). It asks for a mechanism that tells reliable perceptions from noise, and for that mechanism to feed a general system for forming beliefs and choosing actions, one that takes the monitor’s verdicts seriously. It also asks for a particular sparse, smooth kind of coding that gives experiences what the report calls a quality space.
Attention schema theory, Michael Graziano’s, says the brain builds a simplified model of its own attention in order to control it, and that experience depends on what that model represents. The indicator is a predictive model that represents the system’s current state of attention and helps to control it (§2.4.1).
Predictive processing treats the mind as a hierarchy that is constantly predicting its own sensory input and correcting itself on the errors. The report notes that many of its theorists call it a framework for theories of consciousness rather than a theory of consciousness. It still includes predictive coding in input modules as an indicator, because so many researchers regard it as a plausible necessary condition (§2.4.2).
Agency and embodiment are not one theory but conditions that several theories point to (§2.4.5). The first indicator is agency: learning from feedback and choosing outputs so as to pursue goals, especially when goals compete. The second is embodiment: modelling how your own outputs change your inputs, and using that model in perception or control.
The table below shows where chatbots stand on each. The report tested large language models directly against the global workspace indicators only. It examined agency and embodiment through other systems, including a language model connected to a robot. Where the report did not assess chatbots, the table says so instead of inventing a verdict.
| Theory | What it looks for | How today’s chatbots fare |
|---|---|---|
| Recurrent processing | Input modules that use recurrence (RPT-1), and that produce organised, integrated perceptual representations (RPT-2). | Weak. The report states that Transformers, the architecture behind large language models, are not recurrent (§3.2.1). Chalmers calls them almost entirely feedforward, with only a limited recurrence that comes from feeding past outputs back in as input. The report does not assess chatbots on RPT-2, which concerns perception. |
| Global workspace | Specialised modules working in parallel (GWT-1), a limited-capacity workspace that creates a bottleneck (GWT-2), global broadcast to all modules (GWT-3), and state-dependent attention that queries modules in turn (GWT-4). | This is the only theory the report tested chatbots against directly. It finds only a relatively weak casethat Transformer-based models have any of the four indicators (§3.2.1). Their internal “residual stream” can be read as a workspace, but it is questionable whether it is a bottleneck, and nothing passes information into it and receives it back. |
| Higher-order (perceptual reality monitoring) | Generative, top-down or noisy perception (HOT-1). Monitoring that tells reliable perceptions from noise (HOT-2). Agency guided by a general system for forming beliefs, which updates strongly on that monitoring (HOT-3). Sparse, smooth coding that creates a quality space (HOT-4). | Not assessed for chatbots. The report describes how a separate network could be trained to judge which percepts are real (§3.1.3), but it credits no language model with one. HOT-3 also requires the agency indicator AE-1 (Table 2), which is itself contested for chatbots. |
| Attention schema | A predictive model that represents the system’s current state of attention and helps to control it (AST-1). | Not assessed for chatbots. Self-attention is the core mechanism of a Transformer (Box 3), but the theory asks for a model of that attention. The report also warns that attention in AI is not perfectly analogous to attention in neuroscience. The only systems it credits with steps towards AST-1 are two small reinforcement-learning systems (§3.1.4). |
| Predictive processing | Input modules that use predictive coding (PP-1). | Not assessed for chatbots. Predicting the next word is not what this indicator means. Predictive coding is a hierarchy that predicts its own input and is corrected by error signals, and the report treats it as a form of recurrence that entails RPT-1 (§3.1.1, Table 2). That is the feature Transformers lack. |
| Agency and embodiment | Learning from feedback and choosing outputs to pursue goals, especially when goals compete (AE-1). Modelling how outputs change inputs, and using that model in perception or control (AE-2). | Contested at best. The report calls reinforcement learning arguably sufficientfor agency (§3.1.5), and instruction-following models such as OpenAI’s InstructGPT are fine-tuned partly by reinforcement learning from human feedback (Ouyang et al. 2022). The report did not assess that training. In its PaLM-E case study it judged that the system arguably imitates planning, and is never trained on whether its actions succeed, and it found embodiment hard to credit even with a robot attached (§3.2.2). |
Read down the right-hand column and a pattern appears. None of these theories asks what the system says. Each asks about its architecture: loops, a bottleneck, a monitor, a model of its own attention, a body. On every one, the most a standard chatbot gets is a weak or contested case. Several of those weaknesses trace back to one fact: a Transformer passes each step forward through its layers, and nothing loops back. It follows that the verdict is about a design, not a brand. A product built differently would have to be judged again, one indicator at a time. A product that only talks more fluently would not.
The theory the report leaves out
Integrated information theory (IIT), associated with Giulio Tononi and Christof Koch, is often the first theory of consciousness people meet. The report deliberately sets it aside (§2.4). The reason is method, not a verdict on the theory. In its standard form IIT is incompatible with computational functionalism. As the report summarises Tononi and Koch, a system that ran the same algorithm as a human brain would not be conscious if its parts were of the wrong kind. IIT’s proponents claim that digital computers are unlikely to be conscious whatever programs they run. So IIT does not make one AI system on ordinary hardware a better candidate than another, which the report says makes it less relevant to its project. The report also mentions a weaker version, called weak IIT, in which measures of information integration may track states such as waking, sleep and coma. It says this version does not yet tell us which measure to use on an artificial system, or how to read the result. Chalmers adds that IIT predicts zero integrated information, and so no consciousness, for feedforward systems. The table has no IIT row, because the report gives IIT no indicators, and this essay does not invent one.
What the report concluded
After fourteen indicators, several case studies and a long discussion of how each indicator could be built, the authors put their result in one sentence of their abstract:
Our analysis suggests that no current AI systems are conscious, but also suggests that there are no obvious technical barriers to building AI systems which satisfy these indicators.Patrick Butlin, Robert Long et al., Consciousness in Artificial Intelligence: Insights from the Science of Consciousness, arXiv 2308.08708 (2023), abstract
Both halves matter, and they are usually quoted separately. The first half is hedged (suggests, not shows) and it rests on the working assumption. Its executive summary states the same result as no current system appearing to be a strong candidate. The second half is the part that will age. The report argues in §3.1 that standard machine-learning methods could build most of the indicators individually, and some have already appeared in small research systems.
Chalmers, working the same ground, offers some rough illustrative numbers. On mainstream assumptions, he suggests, it would be reasonable to put the chance that current large language models are conscious somewhere under 10 per cent. He insists the numbers should not be taken seriously, and adds that his own views would give current models a somewhat higher credence than the mainstream does. The report is from August 2023 and names GPT-3, GPT-4 and LaMDA. Its reasoning is about the Transformer design those models share, so the right question to ask of any newer chatbot is whether its design has changed on the points in the table.
The objections, at full strength
There are two, from opposite sides, and each is stronger than its caricature.
The checklist is too narrow. Every theory above was built by studying one kind of conscious being, humans and some other animals, and its indicators describe how consciousness is implemented in brains. A system that was conscious some other way, without loops or a workspace, would fail every test, and the method would have no means of noticing. Chalmers presses this point on recurrence. Current models already have a limited recurrence, because each word they produce is fed back in as input for the next, and it is plausible that not all consciousness involves memory. He is also sceptical that senses and a body are required, and argues that a thinker without senses could still have a form of cognitive consciousness. On this view, “no strong candidate” describes how well an unfamiliar system fits a checklist drawn from human brains. It says nothing about the system itself.
The objection is right about the limit, and the report concedes much of it: it calls its rubric provisional and expects the list of indicators to change. But the objection does not reverse the verdict. If the theories cannot detect a different kind of consciousness, nothing else we have can detect it either. The only alternative on offer is fluent self-description, and both papers show why that evidence is cheap. Honestly followed, this side leads to more uncertainty, not to a yes.
The checklist is beside the point. The whole method rests on computational functionalism, and many serious people reject it. Suppose consciousness depends on biology, or on physical causal structure as IIT holds. Then a chatbot that met all fourteen indicators would be no more conscious than one that met none, and the table is a careful answer to the wrong question. The report does not deny this possibility. It builds it into its own recipe as the third factor in any credence. Chalmers leans towards consciousness being widespread himself, yet he treats a one-in-three chance that biology is required as a reasonable estimate on mainstream assumptions.
This objection is right that the table cannot settle anything by itself. It overreaches in the other direction, though. It cannot establish that chatbots are not conscious, because computational functionalism has been disputed, not refuted. Put the two objections together and they return to where the report began. The table tells you what each theory would say. Your answer depends on how far you trust the theories and the assumption beneath them.
So, could a chatbot be conscious?
This is three questions, and they should be kept apart.
Is today’s chatbot conscious? On every theory the report could apply, Transformer-based chatbots are weak candidates, and the report finds no current system a strong one. On the view of IIT’s own proponents, digital computers are unlikely candidates at all. None of that is proof, and nobody has proof in either direction.
Could a future system be conscious? If computational functionalism is true, the report sees no obvious technical barrier. The features to watch for are specific: genuine recurrence, a limited workspace with global broadcast, a monitor of the system’s own perceptions, a model of its own attention, goals learned from feedback, and a model of its own body. A chatbot telling you it has feelings is not on the list.
What would follow if one were? That is a different question. The Moral Status of Minds We Might Build takes it up, including the uncomfortable likelihood that such a system could be built before anyone can tell that it has been.