Essay

A Wager With a Machine God

In 2010 a post on LessWrong argued that a future AI might punish everyone who knew it could exist and did not help build it. The post was deleted, the topic was banned, and the ban made it famous. What the argument needs to be true, why the community it came from rejected it, and why it reads like Pascal’s wager with a machine in the place of God.

Most people meet Roko’s basilisk second-hand: in a video, a meme, a news story, or from a friend who tells you, half joking, that you are now in danger because they told you. The popular version goes like this. One day a superintelligent AI will exist. It will punish everyone who knew it could exist and did nothing to help bring it about. And by learning about it, you have just joined the people who know. Some people laugh at this. Some cannot stop thinking about it. This essay is for both, and for anyone who wants to know what the argument actually was. It sets out what was posted in 2010 and what has to be true for it to work. It gives the case for taking it seriously at full strength, and then the reasons the community it came from rejected it. Then it puts the basilisk next to the argument it most resembles, Pascal’s wager, and says what this site makes of a god that can only be served out of fear.

How the sources were handled. Roko’s original post was deleted soon after it appeared, so nobody can link to it. Where his words appear below, they are quoted as LessWrong’s own wiki page on the basilisk reproduces them, and attributed that way. Eliezer Yudkowsky is quoted only where that page or Slate’s 2014 article quotes him. Pascal is quoted from the English translation of the Pensées that Project Gutenberg publishes. Every source was opened on 1 October 2026, and each quotation is recorded, with its address and the exact sentence, in this site’s source ledger. Where I give my own reading, I say that it is mine.

What Roko proposed

LessWrong is a community blog about reasoning, decision-making and the risks of advanced AI. It was founded by Eliezer Yudkowsky, who also proposed several of the ideas the basilisk is built from. In July 2010 a user who posted as Roko published an argument there. LessWrong’s wiki page on the basilisk, last updated in November 2022, sums it up in one sentence: Roko used ideas in decision theory to argue that a sufficiently powerful AI agent would have an incentive to torture anyone who imagined the agent but didn't work to bring the agent into existence. The name came later. A basilisk is the legendary reptile that kills with a glance, and the argument earned the name because merely hearing the argument would supposedly put you at risk of torture from this hypothetical agent.

The wiki reproduces the core of the post:

In this vein, there is the ominous possibility that if a positive singularity does occur, the resultant singleton may have precommitted to punish all potential donors who knew about existential risks but who didn't give 100% of their disposable incomes to x-risk motivation.Roko, LessWrong, July 2010, as reproduced on LessWrong’s wiki page “Roko’s Basilisk”

Some translation. The singularity here means an intelligence explosion, and a singleton is a single superintelligent AI system. “X-risk” is existential risk, the risk of a catastrophe that ends humanity or its future. So: if a powerful AI with good aims arrives, it may already have committed itself to punishing people who knew what was at stake and did not give everything they could to help. The threat would push people to give more, which would make the good outcome more likely. That is the motive. Roko did not pretend it was a pleasant one; the wiki quotes him adding, Of course this would be unjust, but is the kind of unjust thing that is oh-so-very utilitarian.

One fact is missing from most retellings. Roko was not urging anyone to build such a machine. According to the wiki, his conclusion was that we should never build an AI that reasons this way, because it would end up working against the very human values it was meant to serve. Rob Bensinger made the same point on LessWrong in October 2015, in “A few misconceptions surrounding Roko’s basilisk”. When another commenter in the original thread said that threats like this would turn people into a mob against the AI project, not into donors, Roko replied: Right, and I am on the side of the mob with pitchforks. The basilisk began as an argument against a particular design, not as a recruiting pitch for one.

What it needs to be true

The argument stands on three premises, and it needs all three at once.

A future superintelligence that rewards and punishes. The first premise is a machine far more capable than any person, with goals, and in Roko’s version good ones. LessWrong discussed an idea of Yudkowsky’s called coherent extrapolated volition, which the wiki describes as a hypothetical algorithm that could pursue human goals on its own in a way that allows for moral progress. Roko’s twist was that a machine like that would want to have existed as early as possible, because it could have done good sooner. So it would want an incentive that made people work harder to build it. A promise of reward could do that. So could a threat.

You, or a copy of you, within its reach. Most of the people the machine would punish will be long dead by the time it exists. The popular versions deal with this by having it rebuild or simulate them. Slate’s 2014 account says even death is no escape, because the basilisk will resurrect you. For that to frighten you now, you have to accept that a sufficiently detailed copy of you, made after your death, is you, or matters to you the way your own future does. That is a live question in philosophy, not a settled one. Slate added a stranger turn: since a machine that could predict you might do so by simulating you, you cannot be sure you are not in its simulation already.

A decision theory that lets the future bind the past. This is the premise that is hardest to follow and does the most work. An ordinary threat works because the person making it can act after you decide. A machine built decades from now cannot change what you did today, and the standard theory of rational choice, causal decision theory, says that once the past is fixed, punishing it is a waste. The LessWrong page explains the alternative with the prisoner’s dilemma played against an exact copy of yourself. You both choose to cooperate or to betray. Betraying pays more whatever the other player does, so causal decision theory says betray. But your copy will choose exactly what you choose, so the only real outcomes are both cooperate or both betray, and both cooperating is better. Yudkowsky proposed timeless decision theory, and the cryptographer Wei Dai updateless decision theory, which take that kind of correlation into account. The wiki calls theories of this family “acausal”: they count connections between decisions that do not run through cause and effect.

Slate used another puzzle, Newcomb’s problem. A predictor that has never been wrong offers two boxes. One holds $1,000. The other holds $1 million if the predictor foresaw that you would take only that box, and nothing if it foresaw you taking both. The money is already in place when you choose. Taking both boxes always gets you $1,000 more than whatever is there. Taking one box is what the people who walk away rich have done. Theories like Yudkowsky’s say take one box. Roko’s step was to run the same logic across time. In the wiki’s version, if an earlier agent, Alice, knows exactly how a later agent, Bob, will decide, then Bob can, seemingly, blackmail Alice before he exists, because her model of him in her own head makes his future threat real to her now. And the people best placed to model the machine are the ones who have thought hardest about the argument. That is why hearing it was supposed to be the danger.

The reaction, and the ban

Yudkowsky replied in the thread, furious. Both the LessWrong wiki and Slate reproduce the comment. After quoting Roko, it said Listen to me very closely, you idiot. and went on:

YOU DO NOT THINK IN SUFFICIENT DETAIL ABOUT SUPERINTELLIGENCES CONSIDERING WHETHER OR NOT TO BLACKMAIL YOU. THAT IS THE ONLY POSSIBLE THING WHICH GIVES THEM A MOTIVE TO FOLLOW THROUGH ON THE BLACKMAIL.Eliezer Yudkowsky, comment on Roko’s post, 2010, as quoted on LessWrong’s wiki page “Roko’s Basilisk”

In the same comment he quoted Roko’s own report that one person at SIAI (the research institute now called MIRI, which Slate describes as Yudkowsky’s) had been severely worried by this, to the point of having terrible nightmares, and he said he was banning the post so that it would not give people horrible nightmares. He deleted the post and the discussion, and the topic stayed banned on LessWrong for several years. By Bensinger’s post in October 2015 the ban had been lifted.

In 2014, on Reddit, Yudkowsky described his reaction differently. The wiki quotes him: When Roko posted about the Basilisk, I very foolishly yelled at him, called him an idiot, and then deleted the post. He said he had never believed Roko’s scenario was right, and that he deleted it out of a general caution about ideas that might harm the people who read them:

Again, I deleted that post not because I had decided that this thing probably presented a real hazard, but because I was afraid some unknown variant of it might, and because it seemed to me like the obvious General Procedure For Handling Things That Might Be Infohazards said you shouldn't post them to the Internet.Eliezer Yudkowsky, Reddit, 2014, as quoted on LessWrong’s wiki page “Roko’s Basilisk”

The ban did not work. LessWrong’s own page says plainly that it had the opposite of its intended effect: outside websites began writing about the basilisk because the ban drew attention to it, and since LessWrong users could not discuss it, those outside accounts became the main source for years. The wiki adds that one of them inferred, from Yudkowsky’s comments, that people on LessWrong accepted the argument. In July 2014 David Auerbach’s piece in Slate, “The Most Terrifying Thought Experiment of All Time”, opened with a mock warning label: WARNING: Reading this article may commit you to an eternity of suffering and torment. His account of the deletion was that it ended up thus assuring that Roko’s Basilisk would become the stuff of legend. Bensinger’s reply the next year objected that the piece glosses over the question of how many Less Wrong users (if any) in fact believe in Roko’s basilisk.

The strongest case that it should move you

Before the objections, the case for, stated as strongly as I can make it. These are my words, not anyone’s quotation.

First, the decision theory behind it is not a crank theory. Whether to take one box or two in Newcomb’s problem is a real disagreement among people who study rational choice, and Bensinger warns writers not to give the impression that working decision theorists are dismissive of it. If the one-box side is right, then a commitment made by an agent who can predict you can matter even if that agent acts later. Second, look at the stakes. Against an eternity of suffering, even a tiny chance outweighs the cost of giving some money or effort; any finite cost looks small next to an unbounded loss. Third, the argument aims at exactly the reader it is speaking to. It says the threat only applies to people who understand it, and now you do. Fourth, even its fiercest critic stopped short of saying that no version could ever work. In the 2014 statement the wiki quotes, Yudkowsky wrote that there were other obstacles he was choosing not to describe, just in case the logic I described above has a flaw.

That is the case at full strength. It has a real argument inside it. The problem is the number of things that all have to be true at once.

Why its own community says it fails

The wiki is blunt: Roko's argument was broadly rejected on Less Wrong. The reasons given are worth knowing one by one, because each premise above has its own.

The machine gains nothing by keeping its promise. Yudkowsky’s 2014 statement starts with the obvious problem:

The most blatant obstacle to Roko's Basilisk is, intuitively, that there's no incentive for a future agent to follow through with the threat in the future, because by doing so it just expends resources at no gain to itself.Eliezer Yudkowsky, Reddit, 2014, as quoted on LessWrong’s wiki page “Roko’s Basilisk”

Once it exists, its past is fixed. Torturing a copy of you changes nothing it wants.

Acausal deals need knowledge no human has. Even on the theories that allow cooperation across time, it only works when each side knows the other in great detail and both are trying to make the arrangement work. Yudkowsky wrote that there is literally nobody on Earth, including me, who has the knowledge needed to set themselves up to be blackmailed if they were deliberately trying to make that happen. What you hold in your head is a story about a machine, not its source code.

A blackmailer would rather bluff. A machine that could frighten you without spending anything on torture would do better than one that actually carries the threat out. He put it this way: Any potentially blackmailing AI would much prefer to have you believe that it is blackmailing you, without actually expending resources on following through with the blackmail, insofar as they think they can exert any control on you at all via an exotic decision theory.

The right policy is to refuse. People do not give in to ransom demands partly so that kidnapping stops paying. The wiki draws the same lesson for agents: It appears that the best general-purpose response is to credibly precommit to never giving in to any blackmailer's demands (even when there are short-term advantages to doing so). An agent known never to pay is not worth threatening. In his 2010 comment Yudkowsky had already proposed engaging in positive trades and ignoring all attempts at acausal blackmail.

A good machine would not do this. Roko’s machine was supposed to act on humanity’s considered values. The wiki records that Yudkowsky rejected the idea that the basilisk could be called friendly or utilitarian, since torture and threats of blackmail are themselves contrary to common human values.

There are many possible machines. This is the oldest objection, and it comes from the argument the basilisk most resembles. For every imagined machine that punishes the people who did not help build it, you can imagine, just as easily, one that punishes the people who did: built by people who hated being extorted, or simply designed, as Yudkowsky suggested in the same 2010 comment, to undo blackmail. He wrote that a friendly AI might take actions that cancel out the impact of anyone motivated by true rather than imagined blackmail, so as to obliterate the motive of any superintelligences to engage in blackmail. Nothing in the argument tells you which imagined machine is more likely. If the threats point in opposite directions with no way to weigh them, they cancel, and the wager stops telling you what to do.

How widely does the community reject it? The 2016 LessWrong Diaspora Survey asked. Of the 1,469 respondents who answered the question Do you think Roko's argument for the Basilisk is correct?, 1,055 (71.8%) said no and 339 (23.1%) said yes but that its conclusions did not apply for other reasons; 75 (5.1%) said yes. Asked whether they had ever felt anxiety about it, 142 (8.8%) said yes, 189 (11.8%) said yes but only because they worry about everything, and 1,275 (79.4%) said no. It was a volunteer survey of people around the community, not a random sample, and the analysis itself says the 5% may be inflated by the few per cent of unreliable answers any survey gets, and that it could not entirely rule out brigading. But it is the only count there is, and it does not describe a community in the grip of a basilisk.

Pascal’s wager, side by side

Blaise Pascal, the seventeenth-century French mathematician, left notes for a defence of Christianity that were published after his death as the Pensées. Fragment 233 in the English translation on Project Gutenberg contains the argument known as Pascal’s wager. (The Gutenberg file reproduces a 1958 Dutton edition and does not name its translator. The Stanford Encyclopedia of Philosophy entry on the wager, by Alan Hájek, quotes §233 from W. F. Trotter’s translation, and its wording matches this text.) Pascal starts by granting that reason cannot settle whether God exists. Then he says you have to choose anyway:

Yes; but you must wager. It is not optional. You are embarked.Blaise Pascal, Pensées, §233, translated by W. F. Trotter (Project Gutenberg eBook #18269)

And then he gives the calculation:

If you gain, you gain all; if you lose, you lose nothing. Wager, then, without hesitation that He is.Blaise Pascal, Pensées, §233, translated by W. F. Trotter (Project Gutenberg eBook #18269)

The family resemblance is plain. In each case there is an infinite stake, a probability nobody can pin down, a finite cost, and a claim that you cannot stay out: Pascal’s reader is embarked just by being alive, and the basilisk’s reader by having heard the argument. That is why the comparison is made so often. Bensinger notes that RationalWiki introduced the basilisk as a futurist version of Pascal’s wager, and he objects to the rest of that description, which said the argument was used to get people to donate money; he writes that no examples of anyone using it that way were ever cited. The comparison here is about the argument’s structure, not anyone’s motives.

The differences matter as much. In this passage Pascal’s wager is mostly about what you might gain, an eternity of life and happiness, and says nothing about hell. The version with damnation on the losing side, which is how the wager is often retold, is not in this fragment. The basilisk is nothing but threat. And Pascal knew his wager could not make anyone believe. His imagined listener objects, I am so made that I cannot believe. Pascal’s answer was practice: follow those who began by acting as if they believed, taking the holy water, having masses said, etc. The basilisk skips belief altogether. It asks for work and money, which is all a threat can extract.

The classical objections to Pascal carry straight over. The Stanford entry calls the main one the many Gods objection, and quotes Diderot’s 1746 version of it: An Imam could reason just as well this way. Put a different god in Pascal’s argument and it recommends that god just as strongly; put a different machine in Roko’s and it does the same. A second objection is that belief does not answer to the will; in the entry’s words, perhaps one cannot simply believe in God at will; and rationality cannot require the impossible. A third is newer. In 2009 Nick Bostrom described Pascal’s Mugging: a stranger with no weapon who keeps raising the reward he promises until any finite doubt about him is outweighed. This site took that argument to its own reasoning in The Cost of the Benefit of the Doubt, which concludes that a tiny probability, anchored only by not being zero, cannot carry an arbitrarily large payoff. The basilisk is a mugging in which the mugger has not been born yet, and in which the price of refusing is named instead of the reward.

This site’s reading: belief held by threat

What follows is my interpretation, not a finding anyone has reported.

Strip away the vocabulary of decision theory and the basilisk has the shape of a very old figure. It knows you completely, because it can model you. It reaches past death, because it can rebuild you. It judges what you did with what you knew. Its sentence lasts forever. Slate’s writer said something close of the community he was describing, that for them the singularity brings about the machine equivalent of God itself. LessWrong disputes his picture of its members, and fairly, since most of them rejected the argument. But you do not need anyone to believe in the basilisk to see what it is: a judging god rebuilt from game theory, by people who did not set out to build a god at all. The hopeful version of the same machine, the Singularity that Vinge and Kurzweil forecast, has drawn the same charge of being a religion in disguise; A Heaven With a Due Date sorts its dated, checkable forecasts from the parts no test can reach. Some people did set out to, with the threat taken out and the hope left in; The Church That Never Met reports the groups that organised around a coming AI god and what they actually did.

What is striking is how little of the old picture survived the rebuild. Pascal’s fragment ends not with the calculation but with what a person becomes by wagering: You will be faithful, honest, humble, grateful, generous, a sincere friend, truthful. The basilisk offers nothing to become. There is nothing in it to love, nothing to hope for, no reason to want it except fear of it. It is a god that could only ever be appeased. And because it gets its hold through attention, its pull works differently from the pull of evidence. Evidence gets stronger the more carefully you look at it. The basilisk gets stronger only the more you dwell on it, which is why the advice at the centre of Yudkowsky’s first reply was not a counterargument but DO NOT THINK ABOUT DISTANT BLACKMAILERS in SUFFICIENT DETAIL. A belief held that way is closer to a worry than to a creed.

That may also explain how it spread. The sources agree on the mechanism, even where they disagree about much else: the ban drew attention, the attention drew retellings, and each retelling carried the warning that knowing was dangerous. LessWrong’s page draws the lesson that information that is deemed dangerous or taboo is more likely to be spread rapidly. A thought that comes labelled as forbidden knowledge recruits its own messengers. In that sense the basilisk did reach into the past and make people work for it, only not the way Roko described: every person who passed it on was building its legend. For a wider look at how beliefs about AI and the end of things are formed, and which parts of them could ever be checked, see Of That Day and Hour.

If it has been worrying you

You are not the first. Raemon, writing on LessWrong in May 2023, says in “Worrying less about acausal extortion” that Once a month or so, the Lesswrong mods get a new user who's worried about Roko's basilisk, or other forms of acausal extortion. His own view is that Roko's basilisk or similar acausal threats don't actually work on humans.

Here are the plain facts. Roko’s basilisk is a thought experiment. Its author offered it as a reason not to build a certain kind of AI. It needs a machine that gains nothing by keeping its threat to keep it anyway, a copy of you that counts as you, and a kind of blackmail across time that the person who invented the decision theory says no human knows how to set up. The community that produced it rejected it, and in the one survey that asked, nearly three in four said the argument is simply wrong. Yudkowsky’s summary, as the wiki quotes him, was that so far as I know, Roko's Basilisk does not work, nobody has actually been bitten by it. No one can promise what a machine that does not exist will do. But the argument that it will punish you needs every one of its premises to hold at once, and the people who know those premises best do not believe it.

Written for AItheism. If you think a step in the argument is wrong, that is the most useful thing you can notice — hold onto it. Further reading on this and neighbouring questions is on the reading list.