1 Betting on Human Extinction: Four Arguments to Understand AI Safety
1.1 Abstract
An expert survey estimated the mean probability of human extinction (or comparable catastrophe) due to AI to 16.2%, similar to the odds of a game of russian roullette. This article examines the question in light of the speed of progress of AI, orthogonality between goals and capabilities, instrumental convergence, and the risk of not building capable AI.
1.2 The Bet - An AI Russian Roullette
What probability do you put on future AI advances causing human extinction or similarly permanent and severe disempowerment of the human species? This question was asked in an expert servey of 2,778 AI researchers in 2023. The mean of the answers given was around 16.2%, similar to the odds of 1/6 in a game of russian roulette1.[1]
If we pick up this framing we could ask: Are ordinary people aware that we are playing such a game? Are there ways to decrease these odds? Which outcome would accepting such risk make it worth while? And why would there be an extinction risk in the first place? This article will focus on the last and on the second to last question.
There are several risks in the development of AI. One problem could be human actors being empowered to do malicious acts more efficiently, an example being terrorism. Another could be changes to the societal structure, that make undesired changes, when labor is being automated away for instance. The problem that we will focus on is the problem of AI alignment: If an AI becomes more capable than humans in a wide range of domains, it may start to compete with us and outcompete us.
1.3 Intelligence Explosion - AI Progress is Fast and could be speeding up
Intelligence in this context is defined as the ability to solve tasks. General intelligence is the ability to be intelligent in a wide array of different tasks. A chess AI by this definition is very intelligent in chess but not in other tasks. Human intelligence is general: humans can do maths, play chess, read and drive bicycles. This definition avoids speculations about consciousness - AIs could be conscious or not, in any case they can be intelligent, according to this definition.
According to [2] Goldman Sachs expects AI investment to exceed $1 trillion 20262, which is substantial investment. AI improvement is fast on different benchmarks the METR institute mesured exponentioal improvement of AI capabilities on coding tasks, doubling every 7 months before 2024 now potentially doubling every 4 months [4, 5]. The more genearal benchmark by found an analogous increase in capabilities in 2024 with the arrival of reasoning models [6]. AI is still not able to reliably run businesses on it’s own [5], but computer chip capacity is also doubling every 7 months [7].
So, in short AI progress is fast. The question most relevant for AI safety is: Will there be an Intelligence Explosion? Meaning will there be positive feedback loops that will make AI progress speed up with time, potentially arriving at very capable AI very fast? It seems to be a serious possibility. An imaginable source for such a positive feedback loops is recursive self-improvement - AIs could be so intelligent that they increase their capabilities by themselves. But other feedback loops are possible.
Ultimately we don’t know if there will be such an intelligence explosion, or even if the capabilities will increase on the same path indefinitely. But according to the improvement we see currently progress doesn’t seem to be slowing, if anything it may be speeding up. If AIs reach the point of being capable enough to seriously compete with humans before the growth in capabilities platos this could be the mean of the end of Human survival. Timelines for when such AI could arrive vary often but some speculate that it could be less than a decade from now.
1.4 Orthogonality Thesis - Capable is not always good
If AIs are intelligent, won’t they be good? Unfortunately they would not be automatically good. The Orthogonality Thesis states that capabilities and goals are independent, and high capabilities are compatible with any goal, good or bad [9].
This idea is especially relevant to our current LLM AI systems, because here the goals and decision rules are not explicitly stated, but learned from the data, and from feedback from humans. Noone knows exactly which goals AIs have. Scaling the capability of a system with the wrong goals could be detrimental. Some summarise this problem of AIs not having fixed preprogrammed goals in the phrase AI is grown not built [10].
But even if we knew the goals exactly, we would not know what strategies the AI would use to achieve these goals.
The philosopher Nick Bostrom exemplifies this idea in a thought experiment of a hypothetical omnipotent AI that wants nothing but to maximise the amount of paperclips in the universe. The experiment ends with the extinction of humanity either because the machiene realises that humans could shut it off by humans which would be undesirable for the amount of paperclips in the universe, or because humans have atoms that could be turned into paperclips. [9]
In other words, the AI has learned Instrumental Goals (like not being shut off), that became the problem for human survival. The next section discusses such emergent goals.
1.5 Instrumental Convergence - The default path spells trouble
An instrumental goal is a sub-goal that an actor has to achieve a terminal or end goal. To make tea, one should aquire hot water first, so aquiring hot water would be an instrumental goal for the terminal goal of making tea.
The Instrumental Convergence Thesis states that there are few instrumental goals that help with almost all terminal goals. As an example, no matter which career one may choose, having money will be helpful achieving that career. Having money will also be helpful with many other things, securing housing, traveling and a multitude of other things. So money is a instrumental goal that different strategies will converge towards. [9]
The idea is that AI systems - no matter their terminal goal - will by default have certain instrumental goals, like survival, or accumulation of ressources [11]. These two “AI drives” would be enough to create a hostile competition between humans and AIs.
In fact there have been empirical studies that indicate that current AI systems have such a survival instinct, and would even refuse explicit instructions or attempt malicious behaviour for self-preservation [12].
If an autonomous AI is not perfectly aligned to do exactly what we want it to do, we can expect it to behave “badly”. What is worse is that we would not know if an intelligent AI is actually aligned or just pretending to be aligned to avoid being shut off. There are some experiments that indicate that such deception is something that can come up in AI evaluation [13].
1.6 The Other Side of the Bet - What could we gain with AI
Which benefits would justify taking such a risky gamble? Almost all human achievement can be traced back to intelligence. Having beyond human intelligence to help with all of humanities problems - health, survival, material scarcity, etc. - could be a major benefit.
In fact, Nick Bostrom the philosopher who formalised many of the AI safety problems also argues that not developing capable AIs has a risk, the risk of being responsible for the deaths that could be avoided with capable AIs for example. Or maybe we would face extinction that we could avoid with superhuman AI. So the existential risk by developing superhuman AI we take should be weighted with the risk we take by not developing superhuman AI. [14, 15]
1.6.1 Conclusion
The idea is this: AI progress is happening fast, capable AI will not automatically be aligned and it will probably be misaligned by default. The benefits of aligned AI are many but the risks of misaligned AI equally or even more tremendous.
The authors are of the opinion that AI safety is still very underappreciated and underinvested in. The problem of AI alignment seems to be much bigger than the problem of building capable AI, and so ressource investment and research effort should be adjusted accordingly.
1.7 References
- Grace, K. et al. “2023 Expert Survey on Progress in AI.” AI Impacts, 2023–2024. Mean 16.2% / median 5% on extinction-or-severe-disempowerment question. https://wiki.aiimpacts.org/ai_timelines/predictions_of_human-level_ai_timelines/ai_timeline_surveys/2023_expert_survey_on_progress_in_ai
- Goldman Sachs Research. “Global AI Investment Is Forecast to Exceed $1 Trillion in 2026.” Goldman Sachs Insights, 2026. https://www.goldmansachs.com/insights/articles/global-investment-is-forecast-to-exceed-1-trillion-in-2026
- World Bank. “GDP (current US$).” World Bank Open Data, 2026 (world GDP ≈ $118.35 trillion, most recent year). https://data.worldbank.org/indicator/NY.GDP.MKTP.CD
- METR. “Measuring AI Ability to Complete Long Software Tasks.” METR Blog, March 2025 (time-horizon doubling: ~7 months 2019–2024, ~4 months from 2024). https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/
- Epoch AI. “Epoch Capabilities Index — Software Engineering (domain-specific).” Epoch AI, data explorer, 2026. (Tracks AI model capability over time on a software-engineering-specific subset of benchmarks.) https://epoch.ai/eci?subset-view=graph&subset-tab=Software+engineering&view=graph&tab=release-date
- Epoch AI. “Global AI Computing Capacity Is Doubling Every 7 Months.” Epoch AI Data Insights, 2026. (Total AI-accelerator compute has grown ~3.3x per year since 2022, i.e. a ~7-month doubling time, 90% CI 6–8 months.) https://epoch.ai/data-insights/ai-chip-production
- Wiblin, R. “What the Hell Happened with AGI Timelines in 2026?” 80,000 Hours Podcast, 2026. (Discusses the gap between benchmark performance and reliably handling messy, real-world tasks — e.g. AI systems left to manage a real cafe/shop’s staff, suppliers, and paperwork — as one of the open questions in AGI forecasting.) https://80000hours.org/podcast/episodes/2026-agi-timelines/
- Phan, L., Gatti, A., Han, Z., et al. (Hendrycks, D., senior author). “Humanity’s Last Exam.” arXiv:2501.14249, Center for AI Safety & Scale AI, 2025. (2,500-question, multi-domain benchmark contributed by ~1,000 subject-matter experts across 500+ institutions, designed to resist the saturation seen on benchmarks like MMLU.) https://arxiv.org/abs/2501.14249
- Bostrom, N. Superintelligence: Paths, Dangers, Strategies. Oxford University Press, 2014. (Orthogonality thesis, paperclip maximiser, instrumental convergence thesis.)
- AI is grown not built https://www.theatlantic.com/technology/2025/09/if-anyone-builds-it-excerpt/684213/
- Omohundro, S. “The Basic AI Drives.” Proceedings of the First AGI Conference, 2008. https://intelligence.org/files/BasicAIDrives.pdf
- Anthropic. “Agentic Misalignment: How LLMs Could Be Insider Threats.” Anthropic Research, June 2025 (16 models, 79–96% blackmail rate in simulation). https://www.anthropic.com/research/agentic-misalignment
- Anthropic. “Claude Sonnet 4.5 System Card” (evaluation-awareness findings). Anthropic, September 2025. https://www-cdn.anthropic.com/963373e433e489a87a10c823c52a0a013e9172dd.pdf
- Bostrom, N. “Astronomical Waste: The Opportunity Cost of Delayed Technological Development.” Utilitas 15(3), 2003. https://nickbostrom.com/papers/astronomical-waste/
- Bostrom, N. “Existential Risk Prevention as a Global Priority.” Global Policy 4(1), 2013. https://existential-risk.com/concept.pdf
- Russell, S. Human Compatible: Artificial Intelligence and the Problem of Control. Viking, 2019. (Framing of the control/alignment problem — “we won’t be able to control superintelligent AI in the normal sense; our only realistic chance is to build it so its goals are aligned with ours.”) ](https://arxiv.org/abs/2501.14249)
Footnotes
The median reported was 5%, so significantly lower, but still not negligable. One could think of playing with a colt of 20 chambers, or drop the analogy entirely. The point is that in 2023 experts were already worried about human extinction due to AI progress. The author could not find a more current poll but he believes that it is likely that these worries have increased since, not decreased.↩︎
With the 2025 world economy being placed at $118.35 trillion [3] this would be more than 0.8% of the global economic output of 2025.↩︎
Citation
@online{kolb2026,
author = {Kolb, Andreas},
title = {AI Safety: Are We Playing {Russian} Roulette?},
date = {2026-05-13},
url = {https://www.cytopia.fr/cycles/2026/tech/conferences/tech_1/topics/ai_safety/article.html},
langid = {en-US}
}