~/reading
No. 01July 27, 20269 entries

Game theory, Race dynamics & AI safety

  1. paperAI & Society (2016)
    Racing to the Precipice: A Model of Artificial Intelligence Development

    Stuart Armstrong, Nick Bostrom, Carl Shulman

    The founding formal model of AI race dynamics: teams choose a safety level trading off against capability, solved for Nash equilibria under three information regimes. Counterintuitive result — more information about rivals' capabilities increases danger, not less.

    game theoryAI safetyrace dynamics
  2. paperJournal of Conflict Resolution (2024)
    Uncertainty, Information, and Risk in International Technology Races

    Nicholas Emery-Xu, Andrew Park, Robert Trager

    Extends the Armstrong–Bostrom–Shulman race model to competing states, showing that 'decisiveness' (how much a capability lead translates into victory probability) drives danger, and that transparency about rivals' capabilities is not unambiguously risk-reducing.

    game theorypolicyrace dynamics
  3. paperarXiv (2026)
    Why Open Source? A Game-Theoretic Analysis of the AI Race

    Andjela Mladenovic, Aaron Courville, Gauthier Gidel

    Models frontier labs choosing between open- and closed-sourcing as a formal race game, proving equilibrium computation is NP-hard in the discrete case while remaining tractable via mixed-integer programming.

    game theoryopen sourceAI labs
  4. paperSocial Epistemology (2016)
    The Unilateralist's Curse and the Case for a Principle of Conformity

    Nick Bostrom, Thomas Douglas, Anders Sandberg

    An N-agent decision-theoretic model showing that decentralized unilateral action systematically over-produces risky moves, since the agent who acts is likely the one with the most (miscalibrated) optimistic estimate — even absent malice.

    game theoryAI safetycoordination
  5. paperarXiv (2025)
    Strategic Preemption Under Shared Catastrophic Risk: The Suicide Region and the Race to AGI

    David Tan

    A continuous-time preemption game between AGI developers facing a shared catastrophic externality. Derives a 'suicide region' where both players rationally deploy despite negative expected payoff, and proposes deployment liability and prize-sharing as fixes.

    game theoryAGImechanism design
  6. essaySlate Star Codex (2014)
    Meditations on Moloch

    Scott Alexander

    The foundational essay on multipolar traps — competitive dynamics where individually rational actors collectively produce outcomes nobody wants. Supplies the vocabulary (multipolar trap, coordination failure) used across nearly all later writing on AI race dynamics.

    essaycoordination failuresuperintelligence
  7. essayAI Impacts (2022)
    Let's Think About Slowing Down AI

    Katja Grace

    Directly interrogates the 'racing is rational' assumption by modeling actual race scenarios, concluding that under realistic parameters, safety investment often beats speed even in a competitive equilibrium.

    essayrace dynamicsAI safety
  8. essayAI Futures Project (2025)
    AI 2027

    Daniel Kokotajlo, Scott Alexander, Thomas Larsen, Eli Lifland, Romeo Dean

    A detailed, wargamed scenario modeling US–China competitive dynamics colliding with an intelligence explosion at a composite lab — explicitly modeling how race pressure, not technical inevitability, drives under-testing of a self-improving model.

    scenariorace dynamicsrecursive self-improvement
  9. essaysideways-view.com (2018)
    Takeoff Speeds

    Paul Christiano

    Argues for continuous, broadly-distributed capability gains over abrupt single-actor discontinuities — the crux essay that changes the game-theoretic stakes of racing, since a slow takeoff means no single lab's breakthrough yields decisive strategic advantage.

    essaytakeoff dynamicsrecursive self-improvement