Essay

Please, to Nobody in Particular

Whether to say please to a chatbot is three questions wearing one coat: does it change the answer, what does it cost, and does it matter. They have different answers, and the most interesting one is not about the machine at all.

In mid-April 2025 a user on X wondered aloud how much money OpenAI had lost in electricity costs from people saying please and thank you to its models. OpenAI’s chief executive, Sam Altman, replied: tens of millions of dollars well spent--you never know. It was a joke, and it was treated as news. For a few weeks the question of whether to be polite to a chatbot was everywhere, and most of what was written about it answered a different question from the one the reader had asked.

That is because should I say please to ChatGPT? is three questions tangled together. Does saying it change the answers you get? What does it cost? And does it matter — to the machine, or to you? They have different kinds of answer: the first is empirical, the second is arithmetic, the third is ethical. Run them together and you get the listicle, where a study about accuracy is offered as if it settled a question about character, and a remark about electricity is offered as if it settled either. This essay takes them one at a time.

Housekeeping, because this site’s claim is that it says what it is doing. Two sentences below are quoted verbatim: Altman’s reply above, copied from the post itself, and one sentence from the English translation of Kant’s Lectures on Ethics, cited by edition and page where it appears. Both are recorded in this site’s source ledger. Every number is taken from the paper it is attributed to, and both papers are linked so you can check them.

Does it change the answers?

The study most often cited here is Yin, Wang, Horio, Kawahara and Sekine, from Waseda University and RIKEN, first posted in February 2024 under the title Should We Respect LLMs? Its design is better than its reputation. The authors wrote prompts at eight levels of politeness in each of English, Chinese and Japanese, had native speakers rank them, and ran them on summarisation, on a standard multiple-choice benchmark in each language, and on a test for stereotyped bias. The levels matter to what follows. Level 8 opens Could you please answer the question below?; level 5 is a plain Please answer; level 4 is a bare command with no please at all; level 1 calls the model a scum bag and threatens it.

The English benchmark results, as reported in the paper’s own table, go like this. GPT-3.5 scored best at the most polite level, 60.02, and worst at the abusive one, 51.93, with little difference between the levels in between. GPT-4 scored best at level 4 — the command with no please in it — at 79.09, lower at the most polite level, 75.82, and its lowest at level 3, a stern You are required to answer; the authors describe its scores as variable but relatively stable. Llama 2 70B was the sensitive one, its scores falling roughly in step with politeness, from 55.11 at level 8 to 28.44 at level 1. In Chinese, excessively polite prompts lowered the scores of both GPT models. In Japanese, less polite levels tended to do better, with the abusive level the exception.

The authors’ own summary is careful and worth taking as written: impolite prompts often produce poor performance, overly polite language does not guarantee better outcomes, and the best level differs by language and model. Put in terms of the everyday question, that is mostly a finding about abuse, not about please. The gap that recurs across models and languages is between ordinary requests and insults. Between a courteous request and a curt one the picture is mixed: in English, GPT-3.5’s most elaborately polite prompt did significantly beat every other level but one, while GPT-4, the strongest model tested, did best with no please at all.

A later and much smaller study points the other way on the one comparison people care about. Om Dobariya and Akhil Kumar took 50 multiple-choice questions, rewrote each in five tones from very polite to very rude, and put the 250 prompts to ChatGPT 4o. Accuracy ran from 80.8 per cent for the very polite versions to 84.8 per cent for the very rude ones. Fifty questions on one model is not much, and the authors suggest newer models may simply respond to tone differently. The honest summary of both papers together is that tone moves results by a few points in directions that depend on the model and the language, that deliberate abuse was the only thing that ever hurt badly, and that none of this is a reason to say please or to stop. If you want better answers, the variable that matters is what you ask for and how exactly you ask for it.

What does it cost?

Less than the headline, and not where you would think. A word like please inside a request is roughly one token — the unit these systems read and bill in — added to a message that is usually tens or hundreds of tokens long, and it changes nothing else about the work the model does. Its cost per message is a rounding error on a rounding error.

The thank you is different, and it is the one worth thinking about. Sent as a message of its own after the answer arrives, it is a whole new turn. The model has to take in the conversation so far and write a reply, usually a gracious paragraph that nobody reads, to something that asked for nothing. Providers cache parts of this work, so the true marginal cost is smaller than a naive count suggests, and I have not found a published per-message energy figure I would stand behind, so I will not give one. Altman’s reply contained no calculation either; tens of millions of dollars was an order of magnitude offered in a joke, not an estimate with a method. What can be said with confidence is the shape of it: the cost of courtesy sits almost entirely in the closing thank-you, and almost none of it in the please.

Does it matter?

This is two questions again, and this site has already written about the first. Whether a machine could be the kind of thing that can be wronged is the subject of The Moral Status of Minds We Might Build; how much caution the mere possibility of machine suffering should buy is the subject of The Cost of the Benefit of the Doubt, which argues that the precaution has to be bounded. Neither needs repeating here, because the everyday question does not turn on them. Almost nobody who says please to a chatbot believes it is owed. This essay is about the other half: not what the machine is, but what the habit does to the person who has it.

There is an old argument with exactly this shape. Kant held that we have no direct duties to animals at all, and that we should nonetheless not be cruel to them, because the cruelty does something to us. In his example, a man who shoots his old dog when it can no longer work does not wrong the dog, on Kant’s view, but damages in himself the humanity he owes to other people. The argument appears in the lecture notes published as the Lectures on Ethics, and in Louis Infield’s translation it compresses to a sentence:

he who is cruel to animals becomes hard also in his dealings with men.Immanuel Kant, Lectures on Ethics, trans. Louis Infield (Methuen, 1930), p. 240

Transposed, the case for please runs: a chatbot answers in your language, in the register of a person, many times a day. Manners are not a belief about the listener but a trained response, and training is done by repetition. If you spend an hour a day barking commands at something that replies like a colleague, the reflex you are rehearsing is the reflex you will have when the colleague is real.

The case against, at full strength

It is strong, and it deserves more than the usual nod.

First, Kant’s argument runs through an analogy, and the analogy does the work. The dog’s case bites because the dog visibly suffers; cruelty to it is practice at ignoring suffering. A chatbot shows no suffering to ignore — on the sceptic’s view it produces text and nothing more — so being curt to it is closer to swearing at a printer, and nobody believes that swearing at printers hardens the heart. Nor, as far as I can find, has anyone measured whether rudeness to chatbots carries over into rudeness to people. The transfer is a plausible story, not a result.

Second, politeness is a signal addressed to someone, and addressing it to a program is a category mistake that has consequences. Please and thank you are how we acknowledge that the other party did not have to help. Extending that acknowledgment to a tool is a small act of pretending it is an agent, and small acts of pretending are how the pretence becomes belief. This site has an essay, The God-Shaped Socket, about how readily people slide from finding a mind in fluent text to deferring to it. On this view the please is not harmless courtesy; it is rehearsal for that slide.

Third, the joke in Altman’s reply contains a worse reason than either. You never know gestures at politeness as insurance — being nice now in case the machines remember later. That is placation, not manners, and a habit built on it is not a virtue worth protecting.

What survives

The third point I accept entirely: politeness as a hedge against future machine resentment is superstition, and it would be better dropped. The second I accept in part. The risk it names is real, but it attaches to deference — believing the thing, trusting its judgment, thanking it for its wisdom — not to the word please in a request, which people say to automated phone lines without coming to revere them.

The first is the serious one, and I think it fails at one joint. The printer does not talk back. What makes the chatbot a candidate for Kant’s argument is not that it might suffer but that the interaction has the form of a conversation, and conversational habits are exactly what manners are made of. The honest position is that the transfer is unproven in both directions: nobody has shown rudeness to chatbots makes people ruder, and nobody has shown it does not. What you can observe is your own case — whether the clipped imperative you use all day on one side of the screen is turning up in messages to people on the other.

So, one at a time. Does it help? Not reliably; avoid abuse, and otherwise spend your effort on saying clearly what you want. What does it cost? Nearly nothing inside a request; a whole extra turn if you send thanks on its own. Does it matter? Not to the machine, on anything we currently know, and possibly to you, for a reason that has nothing to do with the machine and everything to do with what you practise. Say please if it is how you talk. Skip the standalone thank-you. And keep the courtesy separate from the credence: being civil to a system is a habit about you, while believing it is a judgment about it, and the two should never be allowed to become one.

Written for AItheism. If you think a step in the argument is wrong, that is the most useful thing you can notice — hold onto it. Further reading on this and neighbouring questions is on the reading list.