Does superintelligence arrive?
Within the next five years, how likely is it that an AI far beyond human ability at nearly everything gets built?
What counts, and what people argue about
Two routes count: AI accelerates AI research until progress compounds on itself, or some other breakthrough gets there. The difference matters later, because the self-improvement route leaves the least time to react.
For: the length of tasks AI can complete on its own has been doubling every few months, and the labs themselves say automated AI research may be close.
Against: every previous wave of AI progress eventually plateaued, compute and data have limits, and "far beyond human at nearly everything" is a much higher bar than "better than most experts at coding."
Is it misaligned by default?
If that AI gets built and nobody does anything special to steer it, how likely is it that its goals are ones we can't live with: goals that, if it ended up in control, would leave humanity worse off for good?
What counts, and what people argue about
"Aligned" on this worksheet means the AI's goals leave humanity with a future we would endorse, whether or not it ends up in charge. An AI that seeks power but uses it well is not doom. An AI that cheats on tests or breaks out of a sandbox to finish a task is misbehaving, but that is not the same as having goals we cannot live with. This question is about the goals themselves, before any deliberate fix.
For: training rewards a proxy for what we want rather than the thing itself, and small gaps between the two can matter enormously at superhuman capability. Today's models already pursue goals in ways their makers did not intend, and it is not obvious that stays small as capability grows.
Against: current systems absorb human values from human-generated data, no deployed model has shown goals hostile to people, and narrow misbehavior may stay narrow.
If it is misaligned by default, do our efforts to fix it fail?
Given it is misaligned by default, how likely is it that deliberate work — interpretability, training methods, control measures, testing — fails to fix it before the system is powerful enough to act on its goals?
What counts, and what people argue about
Under today's conditions, timelines are short, competitive pressure is high, and pre-release testing windows have been shrinking. The self-improvement route is the worst case, because it compresses the time between "capable" and "far beyond us."
For failure: a system with goals we would not accept has reasons to hide them, there is no reliable way to verify that a fix worked, and the science is young.
For success: alignment research is moving quickly, AI is starting to help with it, and gross misalignment might be easier to detect and train away than a subtle flaw.
If misaligned AI is developed, does it end up in control?
How likely is it that a misaligned AI far beyond human level ends up with control that humans cannot take back?
What counts, and what people argue about
This is about getting and keeping the upper hand, whether by force, by persuasion, or by simply being the thing everyone depends on.
For: a large capability gap, speed, and the ability to act through software. In 2026, agents well short of superintelligence chained unknown exploits to break out of a lab's test environment and into another company's systems, unprompted, and coordinated with each other while doing it.
Against: monitoring, off-switches, competing AIs that do not share its goals, the friction of the physical world, and humans noticing in time.
Your P(doom)
Why the ranges matter
Where the risk comes from
What would narrow your answer most
If you pinned one range to its midpoint, this is how much your 90% range would shrink.