Game theory, Race dynamics & AI safety
- paperAI & Society (2016)Racing to the Precipice: A Model of Artificial Intelligence Development
Stuart Armstrong, Nick Bostrom, Carl Shulman
The founding formal model of AI race dynamics: teams choose a safety level trading off against capability, solved for Nash equilibria under three information regimes. Counterintuitive result — more information about rivals' capabilities increases danger, not less.
game theoryAI safetyrace dynamics - paperJournal of Conflict Resolution (2024)Uncertainty, Information, and Risk in International Technology Races
Nicholas Emery-Xu, Andrew Park, Robert Trager
Extends the Armstrong–Bostrom–Shulman race model to competing states, showing that 'decisiveness' (how much a capability lead translates into victory probability) drives danger, and that transparency about rivals' capabilities is not unambiguously risk-reducing.
game theorypolicyrace dynamics - paperarXiv (2026)Why Open Source? A Game-Theoretic Analysis of the AI Race
Andjela Mladenovic, Aaron Courville, Gauthier Gidel
Models frontier labs choosing between open- and closed-sourcing as a formal race game, proving equilibrium computation is NP-hard in the discrete case while remaining tractable via mixed-integer programming.
game theoryopen sourceAI labs - paperSocial Epistemology (2016)The Unilateralist's Curse and the Case for a Principle of Conformity
Nick Bostrom, Thomas Douglas, Anders Sandberg
An N-agent decision-theoretic model showing that decentralized unilateral action systematically over-produces risky moves, since the agent who acts is likely the one with the most (miscalibrated) optimistic estimate — even absent malice.
game theoryAI safetycoordination - paperarXiv (2025)Strategic Preemption Under Shared Catastrophic Risk: The Suicide Region and the Race to AGI
David Tan
A continuous-time preemption game between AGI developers facing a shared catastrophic externality. Derives a 'suicide region' where both players rationally deploy despite negative expected payoff, and proposes deployment liability and prize-sharing as fixes.
game theoryAGImechanism design - essaySlate Star Codex (2014)Meditations on Moloch
Scott Alexander
The foundational essay on multipolar traps — competitive dynamics where individually rational actors collectively produce outcomes nobody wants. Supplies the vocabulary (multipolar trap, coordination failure) used across nearly all later writing on AI race dynamics.
essaycoordination failuresuperintelligence - essayAI Impacts (2022)Let's Think About Slowing Down AI
Katja Grace
Directly interrogates the 'racing is rational' assumption by modeling actual race scenarios, concluding that under realistic parameters, safety investment often beats speed even in a competitive equilibrium.
essayrace dynamicsAI safety - essayAI Futures Project (2025)AI 2027
Daniel Kokotajlo, Scott Alexander, Thomas Larsen, Eli Lifland, Romeo Dean
A detailed, wargamed scenario modeling US–China competitive dynamics colliding with an intelligence explosion at a composite lab — explicitly modeling how race pressure, not technical inevitability, drives under-testing of a self-improving model.
scenariorace dynamicsrecursive self-improvement - essaysideways-view.com (2018)Takeoff Speeds
Paul Christiano
Argues for continuous, broadly-distributed capability gains over abrupt single-actor discontinuities — the crux essay that changes the game-theoretic stakes of racing, since a slow takeoff means no single lab's breakthrough yields decisive strategic advantage.
essaytakeoff dynamicsrecursive self-improvement