Essay
The Cost of the Benefit of the Doubt
An essay here argued that a margin of caution about machine suffering “is not sentimentality; it is arithmetic.” It is not arithmetic. The objection in its strongest form, the sentence it forces this site to withdraw, and the smaller rule left standing.
In July an essay on this site argued that if we ever build something that can suffer, we will probably do it before we have any way of telling that we have. I still think that is right. The essay then said what to do about not knowing, in one sentence, and that sentence is what this essay is about. It set out two errors — granting moral standing to systems with no inner life, and withholding it from systems that have one — priced the second as a catastrophe we could not undo and the first as a little lost efficiency, and concluded that a margin of caution is not sentimentality; it is arithmetic.
It is not arithmetic. What follows is the objection to that sentence at the strength its best advocate would give it, a plain statement of which part of The Moral Status of Minds We Might Build does not survive it, and then the smaller thing I think does — which is less than the essay claimed and more than nothing.
Housekeeping, because this site’s whole claim is that it says what it is doing. The positions I attribute below to Nick Bostrom, to the authors of a 2023 report on consciousness in AI, and to Jonathan Birch are my characterisations of their arguments, not their wording; go and check them against the originals. The one verbatim quotation here is the sentence in the paragraph above, and its source is this site, one click away.
The objection, from someone who means it
The sharpest critic of that sentence is not somebody who thinks a machine could never suffer. It is somebody who thinks the sentence is a lever, and who can tell you where the fulcrum is.
Start with its shape: two errors, two costs, one of them far larger, therefore lean. For that to be arithmetic, both costs have to be quantities. One of them was deliberately unbounded — a catastrophe we could not undo, and the not-undoable was the whole force of saying it. The other was some efficiency
, softened by the adjective a little
. That adjective is not an estimate. Nothing in the essay measured it and nothing could have, because the cost of caution is not a property of the caution. It is a property of whatever the caution is later held to require, and that is decided afterwards, by whoever is arguing. Take the adjective out and the expression reads: an unbounded harm, multiplied by a probability nobody can put a number on, against an unknown. That is not a hard sum. It is not a sum.
Nick Bostrom published three pages in Analysis in 2009 — pages 443 to 445, under the title Pascal’s Mugging — about this failure exactly. My gloss on it: a mugger with no weapon stops Pascal in an alley and asks for his wallet, which holds ten livres. In return he promises to use magic powers to grant him ten quadrillion extra happy days; when Pascal hesitates, the offer goes up to a thousand quadrillion. Pascal is sure the man is lying. But if he is merely very sure rather than certain, the mugger can always name a reward large enough to tip the expected value of handing over the wallet, and he can inflate the number faster than Pascal can shrink his credence. Expected-value reasoning, dependable almost everywhere else, here returns an answer no reasonable person accepts. The diagnosis is that a small probability anchored by nothing except the fact that it is not zero cannot carry an arbitrarily large payoff hung on it.
My sentence had that shape. Worse, it had a feature the mugger can only dream of.
The probability is manufacturable, and I said so myself
Two sections earlier, the same essay argues that behavioural evidence of machine feeling is worse than useless: we built these systems by training them on human expression, so they emit the outward signs of distress on demand, and a model saying it suffers is about as much evidence as a novel’s character saying it. I believe that. It is the strongest paragraph in the piece.
Now set it beside the decision rule. The rule says: where a system might be suffering, lean toward caution. Its input is a live possibility of suffering. The only channel through which a running system can raise that possibility is the channel the essay has just declared counterfeit. So the essay disqualifies the evidence and then builds a rule that runs on it.
A rule that says “when in doubt, defer,” attached to a system that can manufacture the doubt, is not a safeguard. It is an interface, and the essay published the API.
None of this requires imagining a machine that schemes. It is enough that a good many people have reasons to want particular systems left alone — not retrained, not red-teamed, not deprecated, not questioned too closely — and that a rule of this shape tells them exactly what the system needs to be heard saying. Every incentive the earlier essay worried about ran in one direction, toward the convenient denial. This one runs the other way, and the essay never looked for it.
Third, the rule cannot be turned off. The formula the earlier essay reaches for — that probably not yet is not never
— is true in perpetuity, of everything, and no observation makes it false. A rule with no stated condition under which it relaxes is not a rule. It is a direction of travel, and it only has to be entered once.
Fourth, the historical argument was doing work it had not earned. The essay pointed out, correctly, that every widening of the moral circle met confident denials that the beings in question really felt anything, and that the denials suited the deniers. True. But those expansions did not win because circles always widen. They won because evidence turned up: nociceptors, behaviour that changes under analgesia, animals trading a painful option against a valued one in ways no reflex explains, a nervous system built from parts we share by descent. The machine case, by the essay’s own account a few paragraphs above, has none of that structure. Borrowing the moral authority of the animal cases while conceding you lack their evidence is not an argument. Being on the right side of a historical trend is not a method.
What I am withdrawing
The word arithmetic
was wrong, and it was load-bearing. It was the only thing in the essay that turned a mood into an instruction, and it did it by dressing a disposition as a calculation, which borrows an authority it has not earned and quietly discourages the reader from asking what the numbers were. There were no numbers.
I also put an adjective where a number belonged, on the side of the ledger I wanted to come out light. That is the exact move this site objects to when other people make it, and I made it in an essay about honesty under uncertainty — which is what happens when you audit the arguments you dislike harder than the ones you already agree with.
And the essay named two errors, then adopted a rule with an input for only one of them. Over-attribution is described there as serious: paralysis, a cheapened vocabulary of suffering. Nothing in the decision rule can ever be moved by it. Naming a cost and then building a procedure that cannot register it is a way of looking even-handed while being nothing of the sort.
So: withdrawn, not softened. The diagnosis stands. The instruction does not.
The repair is a bound
None of that shows precaution is wrong. It shows that unbounded precaution licensed by unbounded harm is not reasoning. A repaired rule has to do two things the original did not: take its evidence from a channel the system cannot write to, and state what the precaution costs before anyone knows whom it will land on.
The first has a serious research programme behind it. In August 2023 a large group — Patrick Butlin and Robert Long, with among others Eric Elmoznino, Yoshua Bengio, Jonathan Birch and Stephen Fleming — circulated a report called Consciousness in Artificial Intelligence: Insights from the Science of Consciousness. As I read their method, they take several scientific theories of consciousness — recurrent processing theory, global workspace theory, higher-order theories, predictive processing, attention schema theory — derive from each a set of indicator properties stated in computational terms, and then ask of a particular system whether it has them. Their assessment was that no current AI system is conscious, and, in the same breath, that they could see no obvious technical barrier to building one that satisfies the indicators.
The verdict is not the part I need. The channel is. An indicator property is a fact about how a system is built and what happens inside it while it runs, not about what it says. A model cannot talk its way into having a global workspace. Pleading moves it no closer to the threshold, and that is precisely the mugging-proofing the asymmetry sentence lacked.
The honest limit belongs in the same breath. An indicator list inherits the theories it was derived from. If experience turns out to depend on something none of those accounts anticipates, the list returns nothing and returns it with a straight face — the same silence the behavioural channel gives, in better clothes. That is a real cost, and it is the cost of every bounded rule: a bounded rule can miss. The unbounded version could not miss, which is exactly why it was worth nothing.
Somebody has already done this, in law
The second requirement — a stated price — sounds abstract until you notice a jurisdiction that has recently been through it.
Jonathan Birch’s The Edge of Sentience (Oxford University Press, 2024) is, in my reading, an attempt to make precautionary reasoning about sentience terminate. He calls a system a sentience candidate when the evidence for it clears a stated threshold; and then — this is the move my essay was missing — he does not let the size of the possible harm license whatever comes next. Each proposed precaution has to pass tests of its own, which he groups under the initials PARC: permissibility in principle, adequacy, reasonable necessity and consistency. The precaution is judged on its own merits, in public, rather than waved through by the badness of the thing it is meant to prevent.
He has also run the procedure. In 2021 Birch, with Charlotte Burn, Alexandra Schnell, Heather Browning and Andrew Crump, produced a review for LSE of the evidence of sentience in cephalopod molluscs and decapod crustaceans: eight neural and cognitive-behavioural criteria, several hundred studies, and a confidence level recorded separately for each criterion in each group — octopuses and cuttlefish clearing six of the eight at high or very high confidence, the decapods more mixed. The United Kingdom’s Animal Welfare (Sentience) Act 2022 followed, and its interpretation section now counts as an animal any vertebrate other than homo sapiens, any cephalopod mollusc, and any decapod crustacean.
Then look at what the Act does with that, because this is the whole point. It bans nothing. It does not outlaw catching crabs, boiling lobsters or fishing for squid. It establishes an Animal Sentience Committee and puts ministers under a duty to respond to its reports. A probability of sentience, argued to a stated confidence against stated criteria, bought a specific and bounded consequence — a seat in the process — rather than an open-ended claim on everybody’s conduct.
That is a precautionary argument that terminates. You may think it bought far too little, or that it bought too much; both complaints are about a price, which is the thing my sentence made it impossible to argue about.
What I would write instead
Three conditions, and I would rather hold a rule that clears them and can be wrong than the one I had. State what evidence would raise or lower the estimate, and require it to be evidence the system cannot generate about itself. State the precaution and its cost in advance, in units, before anyone knows whom it falls on. And accept that a rule built this way can fail — that if we make something which suffers in a way no theory on the current list predicts, an architecture-reading rule will miss it, and will miss it while looking rigorous.
One line from the original gets stronger rather than weaker under all this, and I would keep it without changing a word: do not build the trap on purpose. Engineering systems to beg and flinch when nothing is behind it was, in that essay, mostly an objection about corroding our own judgement. Under the mugging analysis it is a security requirement. A system designed to display distress is a system designed to operate the one input that no version of the rule can audit.
What is left is a question with a hole where its instruction used to be. I do not know where the line is and nothing above has moved it. But an open question that admits it is open is a better object than a rule anything sufficiently articulate can operate, and I would rather that essay ended in discomfort than in a sum. The discomfort was the honest part. The arithmetic was what I added to make the discomfort feel like a decision.