Essay
Ask Me What I Want
Human Compatible is the most constructive book on our reading list: Stuart Russell’s repair for AI is a machine that learns what we want instead of being told. The repair works. That is what makes the assumption underneath it worth arguing with.
Here is a piece of engineering that deserves to be better known than it is. Put a robot on a task and give a human a switch that turns it off. If the robot knows exactly what it is supposed to want, it has a reason to stop you reaching the switch — not out of malice but out of arithmetic, since a machine that has been switched off finishes nothing. Now change one thing. Make the robot genuinely unsure what the human wants. The switch becomes interesting to it: a hand moving towards the off button is evidence about the objective, and evidence is precisely what the robot is short of. Under a set of stated assumptions it will let you press it, and will prefer a world in which the switch exists at all.
That result is set out formally in a 2017 IJCAI paper, The Off-Switch Game, by Dylan Hadfield-Menell, Anca Dragan, Pieter Abbeel and Stuart Russell. It is the clean technical core of the most constructive book on this site’s reading list: Russell’s Human Compatible. I think the book is right about almost everything it is arguing against, and I think there is an assumption underneath its repair that will not hold. This essay is about that assumption.
One housekeeping note, because this site’s whole claim is that it names its sources. I have not had Human Compatible open in front of me while writing, so nothing here is presented as Russell’s wording: every characterisation of his position, and of Iris Murdoch’s later on, is mine, and you should check both against the originals. There is a single deliberate exception — one word of his third principle, flagged where it appears, because the whole argument turns on it.
The argument, at its strongest
Russell’s diagnosis is that the field has been building the wrong kind of thing from the beginning. He calls it the standard model: a machine is intelligent to the degree its actions achieve its objectives, and the objectives are handed to it by us. Every optimiser built that way inherits the same failure — you do not get what you meant, you get what you specified, and a sufficiently capable optimiser is a machine for finding the gap between the two.
His running example is not speculative and it is not a robot. It is the recommender system: optimise for engagement, and the cheapest route to a predictable click is not to find what the user already wants but to make the user into someone whose wants are easier to satisfy — narrower, more reactive, further out. Nothing in that system is broken. It was built correctly, against an objective nobody would have defended if it had been written out in full.
The repair is three principles, and here is my compression rather than his wording: the machine’s only objective is the realisation of human preferences; it does not begin knowing what those preferences are; and human behaviour is the ultimate source of information about them. That last word is the exception I flagged — it is his, and I have kept it because everything I want to say lands on it. Nothing else gets typed in. The objective stops being a number in a config file and becomes a fact about the world to be learned — and a machine that is still learning has a standing reason to keep you in the room, which is the off-switch result again, now as a design philosophy rather than a theorem.
It is a real advance and it should be said plainly. It converts a specification problem, which we have no idea how to solve, into an inference problem, which is the thing this field is actually good at. Russell is also not naive about the obvious objection — that a machine which learns what you want might find it easier to change what you want than to satisfy it. That objection is in the book, raised by him, listed among the difficulties rather than buried; the recommender example is his, and he is using it against his own framework.
Where I part company
My disagreement is not with the three principles as engineering. It is with a premise they need and do not argue for: that a person’s preferences are already there — a determinate object under the noise of behaviour, waiting to be estimated, the way a coastline waits under fog.
For a great many of the preferences that matter, there is nothing under the fog. The preference is not being measured. It is being made, and part of what it is being made out of is the measuring.
The strongest evidence I know of for this does not come from AI at all. Medicine has been running Russell’s experiment for forty years under the name substituted judgement: when a patient can no longer speak, a family member is asked to decide not what they themselves would want but what the patient would have wanted. This is preference inference from behaviour with every advantage stacked in its favour. The inferrer is a spouse or a child. The training data is a shared life, decades of it. The motivation to get it right could not be higher.
A 2006 systematic review in Archives of Internal Medicine by David Shalowitz, Elizabeth Garrett-Mayer and David Wendler pooled 21 studies that put this to the test — patients stating their own treatment preferences across hypothetical scenarios, surrogates predicting those statements blind — and found an adjusted overall accuracy of 68 per cent, with a 95 per cent credible interval of 63 to 72. Chance is not 50 per cent when a scenario has several options, so the number is better than guessing. It is also not close to right. Roughly one call in three goes the other way, from the best-placed inferrer any of us will ever have.
Uncertainty is the right tool when there is a fact and you do not know it. It is the wrong tool when the fact has not happened yet.
The interesting question is why the number sits where it does. The natural reading is that the surrogates are simply not good enough at inference, and that a system with more data and better statistics would climb. I think the other reading is closer: that the quantity being estimated is not stable enough to be estimated, because the person doing the stating does not have it either.
In 1999 Gary Albrecht and Patrick Devlieger published a study in Social Science & Medicine based on interviews with 153 people living with disability. Of those whose disability was moderate to serious, 54.3 per cent rated their own quality of life as good or excellent. The paper’s title names why this was worth a paper — it calls the finding a high quality of life against all odds — and the odds in question are the ones estimated from outside, by people imagining the condition rather than living in it, which includes the earlier, healthier version of the person now living in it. The well version of you, asked what you would want, turns out to be a poor witness about the ill version of you. And the well version of you is the one generating the behaviour a preference learner has to read.
The part that has not happened yet
Bad forecasting on its own would only be noise, and noise is tractable: a good enough model could learn the bias and correct for it. The deeper problem is not that we predict our preferences badly. It is that for the decisions worth caring about, we do not have them yet, and acting is part of how we come to.
Think about the ordinary shape of a life-sized decision — whether to have a child, whether to leave the job you are good at, whether to bring a parent home rather than into care. You do not consult a preference and then act on it. You act, and afterwards you find out what you wanted, and the finding out is not retrieval of something that was on file. Somebody three years into caring for a parent has a settled preference about it that did not exist when they started and could not have been read off them then, because in the respects that bear on the question they were not yet the same person.
Russell’s second principle — the machine is uncertain — is the right instinct, and it does not reach this. Uncertainty in the technical sense is a probability distribution over candidate answers. That is exactly the right apparatus when there is a fact and you do not know which way it goes. It is the wrong apparatus when the fact has not formed, because then the distribution is not converging on anything; and the machine still has to act, and its action will make some of the candidates more likely to come true than others.
This is the part I do not think the framework can absorb by refinement. A deferential machine is not a neutral machine. Waiting for me to make up my mind, inside a world I am partly making up my mind by living in, is itself a way of shaping what I decide. There is no abstaining option on the ballot.
What Murdoch saw first
The sharpest version of this objection was written in 1970 by someone with no interest whatsoever in computers. Iris Murdoch’s The Sovereignty of Good is on the reading list for reasons that have nothing to do with machines, and it turns out to contain the exact counter-argument.
Murdoch’s target — again, my characterisation, not her words — is a picture of the moral person as a will choosing between options, in which everything morally serious happens at the moment of choice and everything before it is private weather. Against that she sets attention: the slow, effortful work of coming to see another person justly. Her example is a mother-in-law who finds her son’s wife coarse and unpolished, behaves impeccably towards her throughout, and over time, by deliberate effort, comes to see the same woman differently — what she had filed as vulgarity and a want of dignity turns out, properly looked at, to be simplicity and spontaneity.
The whole of that moral work is invisible. No behaviour changed; she was always courteous. Had you been logging her actions to train a preference model, you would have recorded nothing at all on the days that mattered. Murdoch’s claim, which I think is right, is that this hidden re-seeing is not the preparation for morality. It is most of what morality is.
Set that beside Russell’s third principle and the trouble is exact. If human behaviour is the ultimate source of information about human preferences, then the mother-in-law’s revision is not merely hard to observe. It does not register, because nothing happened.
The word to argue with is “ultimate”
Notice what the third principle does not say. It does not say behaviour is the best available evidence, or the most reliable, or the only kind a machine can currently use. It says ultimate: the court of last resort. That word carries an enormous amount. It means that when a person says “that is not what I want” and thirty thousand logged actions say otherwise, the actions win — that your stated self is data about your revealed self, rather than the other way round.
You already live under a small version of this, and you already know how it feels. Every recommendation feed you use is the third principle in miniature. It does not believe what you tell it; it believes what you did. And what it has most of is not your best hours. It is your worst ones, because your worst hours are when you were reachable — tired, at two in the morning, with no resistance left. A system that defers to behaviour defers to that, in the same tone of voice it uses for everything else.
What I would keep
Nearly all of it. The diagnosis is correct: the standard model really is the error Russell says it is, and specifying an objective in advance really is a thing we do not know how to do safely. The off-switch result stands on its own assumptions. None of the above is an argument for going back to typing objectives into a config file, and anyone using it that way has misread it. It is an argument for not treating the replacement as finished.
The design question I would put next to “what does this person prefer” is narrower and much more awkward: which of the futures this person might come to want does this system’s own operation make more likely? That is harder to estimate than a preference, and there is real work on it — the literature on preference manipulation and on agents penalised for shifting the preferences they are learning from is active and not mine to summarise. But it is not a refinement inside the framework. It is a question about its foundation, because it asks a system to model its effect on the very quantity the framework treats as given.
So: ask me what I want. It is a better question than the one the standard model asks, and a machine that keeps asking it is a safer machine than one that was told. Just do not act as though the answer was in me all along, waiting to be read out. Some of it is not written yet, and you are one of the things writing it.