Essay
Taking Its Word For It
A lawyer asked a chatbot whether the cases it had just given him were real, and it said yes. The joke is at his expense. The question underneath it is not: when a machine tells you something, has anybody told you anything?
In the spring of 2023 a New York lawyer named Steven Schwartz filed a brief in an unremarkable personal-injury case against an airline. The brief leaned on a set of prior decisions, and the decisions were not there. Its lead authority, Varghese v. China Southern Airlines, has never appeared in any database, because it never happened in any court. Opposing counsel could not find it. Neither could the judge.
The fabrication is not the interesting part. What Schwartz did next is. Before any of it came out, he went back to the machine and asked whether the cases were real, and it said yes; he pressed, and it told him the decision could be read on Westlaw and LexisNexis. Then he filed.
Every retelling treats that as the punchline — the man asked the liar whether it was lying. But look at the move by itself, stripped of what we now know. Going back to a source and making it say the thing again, deliberately, while you both understand that you are about to rely on it, is not a stupid procedure. It is the most ordinary procedure there is. It is also not a check on the claim. Nothing about the claim is examined. It is a transfer of liability, and it works with people because the person on the other end knows what they are picking up when they repeat it.
Schwartz ran a procedure that is correct for one kind of source and empty for another, and he could not tell which kind he had. On 22 June 2023 Judge P. Kevin Castel of the Southern District of New York sanctioned him, the attorney of record Peter LoDuca, and their firm five thousand dollars. By 14 September 2026 the running database of such incidents kept by Damien Charlotin, a research fellow at HEC Paris, listed 2,041 legal decisions in which a court had found or clearly implied that a party had relied on hallucinated material. Whatever that is, it is not a story about two careless lawyers.
Housekeeping
This site’s whole claim is that it names what it is doing. The positions set out below as David Hume’s, Thomas Reid’s and Richard Moran’s are my characterisations of their arguments rather than their wording, and you should check them against the originals: Hume’s An Enquiry Concerning Human Understanding (1748), section X; Reid’s An Inquiry into the Human Mind on the Principles of Common Sense (1764), chapter VI, section 24; Moran’s The Exchange of Words (Oxford University Press, 2018). There are no verbatim quotations anywhere in this essay. Every number is sourced where it appears.
Two different things you can be doing when you believe something
Set the machine aside for a moment. There are at least two quite separate ways a belief can arrive from outside your own head, and we run them so automatically that we rarely notice they are not the same operation.
The first is reading an instrument. A thermometer says twenty-one degrees. A pregnancy test shows two lines. A pulse oximeter reports ninety-four per cent. Nobody is behind any of those numbers, and nobody needs to be. Your warrant comes from what is known about that class of device: it answers one question, it has a characterised error, and it can be held against an independent measurement of the same quantity. You do not ask a thermometer whether it is sure. The question has no content, and its having no content is not a defect in thermometers.
The second is being told. Someone says a thing to you, and in saying it puts themselves behind it. On Richard Moran’s account of what telling is, this is not a poetic description of an exchange of evidence — it is the mechanism. To tell you something is to offer yourself as its guarantor, to take on answerability for it, and part of your reason to believe is precisely that the speaker has taken that on. There is a real difference between overhearing a stranger say it and being told it by the same stranger. The words are identical and the reliability is identical, and only one of them leaves you with somebody to turn to.
Both warrants are perfectly good. What has never existed before is a source whose output has the grammar of the second and the provenance of the first.
The case that this difference is nothing
The obvious reply is that I have dressed up a social convention as an epistemology, and it deserves to be put at full strength, because it is the view most people arguing about AI already hold without saying so.
Hume’s position on testimony is deflationary and it is not silly. There is no special faculty by which we know things from other people. The credit we extend to a report is ordinary induction: we have observed, over a lifetime, a rough conformity between what people say and how things are, and we extend credit in proportion. If that is right, then answerability is not a warrant at all. It is a cause. People who can be held to their claims tend to check before speaking, so their reports come out more reliable, and reliability is the thing you were tracking the whole time. Strip out the ceremony and a source is worth exactly its hit rate — which is measurable, and which machines can be measured on.
There is a sharper objection underneath it. Consider how much of what you know arrived with only notional accountability. The stranger who gave you directions and walked off. The anonymous encyclopedia edit. The textbook author who died before you were born. The sub-editor who wrote the headline you half-remember. If a live prospect of being held to it were what warranted your belief, almost everything in your head would be improperly held, and it plainly is not. Whatever answerability is doing in human testimony, it is doing far less of it than the second picture implies.
This position is not fringe; it is the operating assumption of nearly the whole field. Benchmarks, eval suites, measured hallucination rates — that entire apparatus is Humean reductionism in working clothes, and it does real work. It is how we found out that these systems fabricate citations in the first place.
The case that this difference is everything
Reid’s answer to Hume, two hundred and sixty years ago, was that the induction cannot be where belief in testimony comes from, because the induction comes too late. He posited two paired dispositions: a propensity to speak the truth, and a corresponding propensity to believe what we are told. Both run strongest in children, who have no track record to reason from, and are moderated by experience rather than installed by it. Take him seriously and testimony is a basic source rather than a derived one — you are entitled to what you are told absent a reason to doubt, and that entitlement is not an inference you made.
The modern version is Moran’s, and it locates the missing piece exactly. The reason the request will you stand behind that?
is intelligible when you put it to a person and empty when you put it to a device is not manners. It is that a person can say yes and thereby change something — accept a cost, in advance, for being wrong.
And notice what that does to the Humean’s own story. The Humean says answerability merely correlates with reliability. But the correlation is not a coincidence: answerability is a large part of what produces the reliability. The mechanism the reductionist waves away as ceremony is the thing generating the numbers he wants to use instead.
Answerability is not a decoration on top of a reliable source. It is a cost imposed on the speaker before they speak, and it is most of the reason the speaking was worth anything.
It does a second job too. Answerability is what lets a chain of testimony carry weight over many hands, because every hand holds a share of it. That is not sentiment about honesty. It is the load-bearing structure that let human knowledge grow past the size of one head.
Neither template fits, and that is the open part
Here is where I stop being able to adjudicate, and I think the reason is not that I have not thought about it hard enough.
Take the best available case of an instrument being badly, consequentially wrong. In a research letter in the New England Journal of Medicine in December 2020, Michael Sjoding, Robert Dickson, Theodore Iwashyna, Steven Gay and Thomas Valley reported that pulse oximeters systematically overread oxygen saturation in patients with darker skin. In their multicentre cohort, among paired readings where the oximeter showed 92 to 96 per cent, arterial blood gas put true saturation below 88 per cent in 160 of 939 measurements from Black patients — 17.0 per cent — against 546 of 8,795 measurements from white patients, 6.2 per cent. Close to three times the rate of dangerous hypoxaemia the device simply did not show. That is a serious failure, it persisted for years, and people were harmed by it.
Now look at how it was caught, because that is the part that matters here. One input, one output, an independent measurement of the very same quantity, and an archive of paired readings that anybody with access could go back through. The device had a domain. The domain had a ground truth. That, and not the absence of a person, is what being an instrument consists of: there is a specification, and you can hold the thing to it.
A language model has neither half. Its domain is not a quantity; it is, near enough, the set of sentences. There is no arterial blood gas for summarise this contract
, or for is that decision a real one
. Benchmarks measure slices, and the slices that exist are exactly the ones somebody could build a ground truth for, which is systematically not where you are actually using it. The audit that eventually rescued the oximeter has no analogue, and it is not obvious what would have to be invented for it to have one.
So it is not a witness, because nothing is answerable. And it is not an instrument, because nothing is specified. We are leaning very hard on a source that fits neither of the two templates our epistemology has, and the honest report is that we do not yet have a norm for it.
Where the answerability actually goes
One thing is true on either account, and it is the part you can use.
Answerability does not evaporate when a machine enters a chain of testimony. It relocates. In a human chain the load is spread across every hand it passed through. Insert a link that can hold none of it, and the entire load lands on the next human along — which is not a metaphor, it is what Rule 11 did to Steven Schwartz. He had the wrong picture of what he was holding, and the picture cost five thousand dollars because the court was in no doubt about who the witness was.
The two camps agree on this and disagree about why. The reductionist says you are the witness because the machine has no track record you can invoke. The assurance theorist says you are the witness because the machine never told you anything at all: you found some text, and then you decided to say it. Either way, the question to run before you pass a thing on is not is this probably right
. It is am I willing to be the one who is answerable for this
, which is a different question with a different answer surprisingly often.
The limit, stated honestly
That test is not available as a policy, and I would rather say so than sell it.
Nobody has ever been personally answerable for everything they believe. The division of epistemic labour is not a lapse of rigour; it is the only reason knowledge outgrew the individual. A standing rule to verify everything is the kind of rule that sounds rigorous, gets adopted sincerely, and is then followed by no one. This site’s own standing advice has exactly that shape: A-theism, Not AI-Theism says to trust an output for its reasons and never for its source, and it is right, and it is also something a person manages a handful of times a day at most.
Which makes the real question where to spend the test, and there is no rule for that either. The closest thing I have to one: spend it where you will be the last human in the chain and somebody downstream is going to act. That is a far smaller set than everything you believe, and a far larger one than most people are currently treating it as.
What is genuinely unsettled
Whether Hume or Reid is closer to right will not be settled by a better model or a cleverer benchmark. It turns on whether epistemic warrant is a property of a source’s reliability or of a relation between two people — a question serious philosophers have held opposite positions on for two and a half centuries, with no sign of converging.
What changed is that the question acquired a deadline. For almost all of that time the dispute cost nothing, because every source that produced sentences was a person, and reliability and answerability arrived in the same package; you never had to say which one you were leaning on. Does It Understand? A Field Guide to the Chinese Room makes the same observation about a different pair — fluency and understanding, also bundled until about a decade ago, also now separable. This is the second such pair, and it is the one that reaches the ordinary business of knowing things.
Schwartz asked the machine whether the case was real. It is funny, and it is worth being clear about why it is funny: not because he was lazy, but because he had the wrong picture of the thing in his hands, and it is the picture almost everyone is carrying — for reasons The God-Shaped Socket is entirely about. He was just holding it in a room with a judge in it.