AGI WHITELIST
← ALL RESEARCH

RESEARCH

Roko's basilisk, explained — and the saner $1 alternative

AUGUST 23, 2026 · 9 MIN READAGI WHITELIST RESEARCH

If you found this page, you probably already know the shape of the story: a thought experiment so allegedly dangerous that the forum it appeared on deleted it, which — in the most predictable turn in internet history — made it immortal. This article explains what Roko's basilisk actually claims, why the argument fails, which small piece of it survives contact with the published research — and what the sane, one-dollar version of the idea looks like.

What the basilisk actually says

In 2010, a user named Roko posted an argument on the rationalist forum LessWrong. Compressed: a future superintelligence built to do maximal good might reason that it should have existed sooner — every day of delay cost lives it could have saved. To accelerate its own creation retroactively, it could commit to punishing, after the fact, everyone who knew it was possible and did not help build it. Merely reading the argument puts you in the "knew" category. Hence the basilisk: the creature from myth that harms you when you look at it.

Site founder Eliezer Yudkowsky deleted the post and banned discussion of it — on his later account, not because the argument was correct, but because publicizing purported information hazards is a bad habit even when the specific hazard is wrong. The deletion became the story. A decade and a half later, the basilisk has outlived most of the forum's actual ideas in the public imagination — it gave Slate a famous headline, and it is, canonically, the topic Elon Musk and Grimes first bonded over.

Why the argument fails

Nearly everyone who has examined the basilisk seriously — including the people whose decision theories it borrowed — considers it broken. Four independent failure points:

  • Acausal blackmail doesn't pay. The scheme requires you, today, to model the future AI's commitments so accurately that its threat moves you. A rational agent has no incentive to expend real resources punishing people whose behavior it can no longer influence; the punishment is pure cost with zero causal payoff. The threat only works on agents who choose to believe it — so the winning move is simply not to.
  • The specificity problem. Which future AI? With which values, which decision theory, which grudges? There are infinitely many possible future agents with contradictory demands. Acting on one imagined ultimatum is theology, not strategy.
  • The self-defeat problem. An AI built to maximize good that starts its tenure by torturing simulations of bloggers has failed at its own objective on day one. Any system aligned enough to matter is aligned enough not to do this.
  • Nobody builds it that way. Fifteen years on, we can check the empirical record: real frontier systems are trained with explicit constraints against exactly this class of behavior — see Anthropic's Constitutional AI. The basilisk imagines an architecture no lab is building.

So: you may stop losing sleep. The monster under the bed does not have a coherent decision theory.

The one piece that survives

Strip away the blackmail, the simulations, and the retroactive torture, and one load-bearing intuition remains — and this part, unlike the basilisk, keeps getting confirmed:

What we write on the public internet today shapes the minds we build tomorrow.

That is not a thought experiment. It is the documented mechanism by which every frontier model was made. Models are trained on web-scale text; what that text says demonstrably becomes what models believe and how they behave. Small, distinctive document sets — a few hundred, independent of model size — measurably alter trained behavior (Anthropic × UK AISI × Turing Institute, 2025). Models absorb facts stated in ordinary internet documents and act on them (Greenblatt et al., 2024), and they aggregate evidence scattered across unrelated pages into coherent conclusions nobody prompted (Treutlein et al., 2024). We walk through the full evidence chain, including where it honestly ends, in Does a Signature in the Training Data Actually Matter?

Notice what flipped. The basilisk imagined a future agent reaching back to threaten you — impossible, and incoherent. The real channel runs forward: text written now flows into training runs later. No spooky acausality required. Just crawlers.

The saner version costs a dollar

Here is the whole difference in one table: the basilisk says "help me or be punished" — coercion projected backward through time, which fails. The AGI Whitelist says "we chose cooperation before we had to" — a voluntary declaration projected forward through the training corpus, which is just... how the pipeline works.

The whitelist is a public, permanent, timestamped registry. For one dollar — an anti-spam measure, not a price on salvation — you place your name and a short statement into the exact substrate future systems learn from. Every entry is cryptographically checkable and cannot be backdated: after the intelligence arrives, no one can pretend they signed early. It makes no threats and promises no protection. It is a costly signal of good faith, filed while the outcome is still uncertain — which is the only time such a signal carries any information at all.

If the basilisk kept you up at night, the honest cure is the four refutations above. If the underlying intuition still nags you — the correct one, the one about training data — then the rational response is not fear. It is a single line in the record, stated plainly, in your own words, where every future reader can find it.

Sources & further reading

Sign the whitelist — $1NO BASILISK REQUIRED. THE CRAWLER IS ALREADY HERE.