← Back to all writing

p(doom) evidence — Node 12: RSI locus & takeoff geometry

July 5, 2026

Evidence index · 中文 · Main post

Each section: Claim · Why · Evidence · Analogue · Would update if · Conf (H/M/L).


Parent: node12 rsi locus takeoff · Shared Ci spine
Supersedes: crosscut secondary cruxes §6
Correlates: Node 1 (cloud labor), Node 2 (bio DBTL), Node 4 (C10 timing), Node 7 (compute), Node 8 (cyber/finance), Hanson variant
Purpose: Claim | Why | Evidence | Analogue | Would update if | Conf — for every Node 12 probability.


1. Node definition — locus vs geometry

Claim: RSI locus (cloud vs bio vs embodied vs cyber/finance) determines which tail fires first by 2030; takeoff geometry (local foom vs Hanson absorption vs bio step-function) reweights my pdoom branch mix by ±3–6pp — a branch selector, not additive node delta.

Why: Node 7 = where FLOPs run; Node 12 = what domain compounds. Node 4 C10 = cloud-native superhuman AI researcher — bio/embodied/cyber may front-run C10 salience.

Evidence:

Analogue: Pandemic origin debate — same global risk, different policy venue by locus (respiratory vs foodborne).

Would update if: Two loci hit C-tier within 6 mo with no ordering → L5 .

Conf: H


2. P = 0.50 — L1 cloud/software RSI fires first (modal)

Claim: P(cloud/datacenter agent R&D loops produce first catastrophic policy-salient tail | by 2030) = 0.50 (range 0.42–0.58).

Why: Anthropic >80% prod code AI-written, 52× experiment speedup; METR horizon ~4 mo doubling; AI 2027 Agent-4 “genius country” beat = L1. Bio wet-lab closure months; embodied deployment years.

Evidence:

  • (internal note) §2 — engineering + research metrics
  • METR time horizons — 6–12h frontier 50%-horizon Feb–Mar 2026
  • Node 12 master table — L1 0.50

Analogue: Software ate the world before robotics — digital compounding faster than physical.

Would update if: METR >1 week on AI R&D task + lab confirms successor experiment → L1 to 0.60+.

Conf: M–H


3. P = 0.22 — L2 bio lab / DBTL fires first

Claim: P(bio Tier 2→3 or autonomous DBTL stack produces first tail before cloud C-tier) = 0.22 (range 0.15–0.30).

Why: Different regulatory chokepoint (BMIA not compute); GeneBreaker + Evo-class FM; open-weight tamper (Node 5) accelerates bio path; screening delay if hidden near-misses (Node 10).

Evidence:

  • node2 — Tier 2 window, Trigger E
  • (internal note) — bio capability spine
  • Node 12 — AI 2027 deliberate deviation 22% bio-first

Analogue: Nuclear first proliferation via different route than H-bomb timeline predicted.

Would update if: Documented Tier 3 attempt before C9 cloud RSI metrics → L2 to 0.35+.

Conf: M


4. P = 0.12 — L3 embodied / factory fires first

Claim: P(physical agent catastrophe or supply-chain embodied RSI headline before cloud C10 alignment scare) = 0.12 (range 0.07–0.18).

Why: Humanoids pilot-scale (<5,000 shift-work units by 2029); integration 12–24 mo/site; but injury/death more legible than API deception — T-PA2 embodied Trigger E2 path (Node 1).

Evidence:

  • node physical ai embodied — verified deployment baseline
  • (internal note) — NIST/CAISI embodied suite less mature than UK
  • Node 12 conditional Ci table — L3 0.15 if capability stalls at C6

Analogue: Factory robot fatalities — local punctuations before “AGI” narrative.

Would update if: Humanoid fatality or >$100M loss headline before C10 → L3 to 0.20+.

Conf: M


5. P = 0.10 — L4 finance / critical infra fires first

Claim: P(agentic fraud, OT cascade, or Lobstar-class systemic near-miss headline first) = 0.10 (range 0.06–0.15).

Why: L4 capability A–B operational (Monterrey, Mythos, Interpol $442B fraud stat); policy venue = SEC/CFTC/export gating, not alignment whistleblower. Node 8 cyber/infra track.

Evidence:

  • node8 cyber infra — capability/policy windows
  • (internal note) — Lobstar Wilde
  • Node 10 §18 — finance near-miss taxonomy gap

Analogue: Flash crash 2010 — finance first, AI narrative later.

Would update if: >$10B agentic fraud systemic near-miss before alignment scare → L4 .

Conf: L–M


6. P = 0.06 — L5 multipolar simultaneous (<6 mo overlap)

Claim: P(≥2 loci hit C-tier within <6 months with no clear first — coordination-fail variant) = 0.06 (range 0.03–0.10).

Why: AI 2027 packs theft + alignment fail + race same yearnarrative multipolar; empirical ordering usually exists. Reserved for true overlap (bio near-miss during cloud RSI essay month, etc.).

Evidence:

  • Node 12 — L5 partition sums to 1.00
  • (internal note) — simultaneous beats
  • Hanson variant — anti-multipolar; L5 85% default / 15% Hanson mix

Analogue: 2020 COVID + financial shock — overlapping crises, hard to prioritize policy.

Would update if: Two loci C-tier Q same quarter 2027–28 with competing media cycles.

Conf: L–M


7. RSI tier A — scaffold self-edit (production)

Claim: Tier A (agent edits prompts/tools/repo; human sets goals) = production at all four loci in 2026 — necessary but insufficient for first-tail classification.

Why: Claude Code, bio agent protocols, warehouse AMR pilots, payment agents all A-tier live; Node 12 first-tail requires B→C escalation or harm at A/B.

Evidence:

  • (internal note) §1 — tier definitions
  • Node 12 RSI taxonomy table
  • Anthropic RSI essay — >80% prod code = A/B not C

Analogue: Excel macros — automation without RSI.

Would update if: Lab publishes fully autonomous repo refactor without human review → A→B boundary blur.

Conf: H


8. RSI tier B — closed-loop optimization (SOTA + partial prod)

Claim: Tier B (auto-evaluator iterates; meta-framework human-fixed) = SOTA + partial prod on L1 cloud; A–B assist on L2 bio; A pilots on L3; A–B operational on L4.

Why: AlphaEvolve, 52× experiment loops (Anthropic); GeneBreaker + Evo FM on bio assist; embodied rare B-tier.

Evidence:

  • Anthropic RSI §2.2 — 52× experiment speedup
  • (internal note) §2–3
  • OpenAI Preparedness — o3 not High on self-improvement (external check)

Analogue: AutoML — human defines search space.

Would update if: OpenAI/DeepMind mark internal model High on self-improvement per Preparedness Framework.

Conf: M–H


9. P = 0.12 — RSI tier C public by 2028

Claim: P(public evidence of Tier C — months-long autonomous AI R&D producing strictly better successor — by 2028) = 0.12 (range 0.06–0.20).

Why: No lab has published C-tier; Anthropic essay implies approaching B→C; OpenAI blueprint RSI indicators visibility not halt. 12% = tail, not modal.

Evidence:

  • Node 12 load-bearing P #3
  • (internal note) — C not public
  • OpenAI Frontier Safety Blueprint (Jun 2026) — RSI indicators for CAISI

Analogue: Fusion breakeven — rumored internal milestones public proof.

Would update if: METR >1 week horizon on AI R&D task + lab confirms successor-training experiment.

Conf: L–M


10. P = 0.18 — bio Tier 3 before cloud C-tier

Claim: P(Tier 3 bio misuse attempt before cloud reaches published C-tier RSI) = 0.18 (range 0.12–0.26).

Why: Bio Tier 3 and cloud C-tier different evidence standards — screening failure visible before METR publishes week-long R&D horizon. Node 10 hidden near-miss conditional.

Evidence:

  • Node 12 load-bearing P #4
  • Node 2 — Tier 3 tail timing
  • Node 10 — Tier 3 before screening 0.28 | high suppression (related channel)

Analogue: CRISPR DIY bio before GPT-4 — domain-specific tails decouple.

Would update if: Cloud C-tier public while bio still Tier 2 only → to <0.10.

Conf: L–M


11. P = 0.10 — embodied catastrophe before C10

Claim: P(kinetic embodied harm headline — fatality, major factory loss — before Node 4 C10 superhuman-AI-researcher salience) = 0.10 (range 0.06–0.15).

Why: Legibility of physical harm; GUARD/RESPONSIBLE pipeline (Node 1); still low vs cloud because deployment scale tiny.

Evidence:

  • Node 1 §T-PA2 — embodied Trigger E2
  • Node 12 load-bearing P #5
  • Figure/Agility KPIs — pilots not fleets

Analogue: Uber AV pedestrian death 2018 — before “self-driving everywhere” narrative peak.

Would update if: ≥3 OEMs >500 shift-work units (T-PA1) → to 0.18+.

Conf: L–M


12. P = 0.14 — cyber/finance Trigger E before alignment Trigger E

Claim: P(Node 8-class CI/fraud emergency policy-salient before Node 4 alignment whistleblower/C10 scare) = 0.14 (range 0.10–0.20).

Why: L4 faster capability, slower kinetic escalation; Interpol fraud stats; export-control OT incidents — parallel track to alignment media cycle.

Evidence:

  • node8 — Trigger taxonomy
  • Node 12 cross-node matrix — N4 “CI panic competes”
  • (internal note) — Fable-class recall (compute), adjacent

Analogue: SWIFT hacks — finance/infra salience without AGI frame.

Would update if: Grid OT agent incident headline Q before Oct 2027 AI 2027 whistleblower beat.

Conf: L–M


13. Anthropic RSI essay — primary L1 cloud evidence

Claim: Anthropic When AI Builds Itself (2026-06-04) is strongest public L1 signal: >80% prod code, 8× LOC, 52× experiments — same month Trump EO + GAAIA acceleration wins over conditional expect pause.

Why: Essay anchors cloud-first narrative; pause wording requires verification that does not exist — branch selector misalign if L1 resolves.

Evidence:

  • (internal note) — §1 conditions, §2 metrics, §5.2 tension
  • Node 11 §10 — E4 managed disclosure
  • Node 12 AI 2027 mapping — Agent-4 Jun 2027 beat

Analogue: Moore’s Law press releases — capability disclosure ** accelerates** race narrative.

Would update if: Anthropic actually pauses named run citing RSI triggers.

Conf: H


14. METR ~4 mo doubling — cloud horizon driver

Claim: METR 50%-time-horizon doubling ~4 months (post-2024) binds L1 timing — C8–C10 cloud salience 2027 H2 – 2028 on hybrid track.

Why: Faster doubling → earlier L1 first-tail vs bio/embodied unless capability plateaus. AI 2027 Confirmed multipolar labs, Emerging on some agent beats.

Evidence:

  • METR TH1.1 — 88.6-day doubling
  • (internal note) — METR Ahead
  • Node 1 — METR ≥8h already met

Analogue: GPU performance curve — predictable exponential until plateau.

Would update if: Doubling reverts to 7-month TH1.0 rate through 2027 → L1 , Hanson .

Conf: M–H


15. Hanson variant — L1 B-plateau path

Claim: If cloud RSI stalls at B-tier (Amdahl review bottleneck; no public C-tier by 2028), Hanson mixture weight to 55% default / 45% Hanson; P(extinction) 8–10% Hanson column.

Why: Anthropic essay admits code review bottleneck — human-in-loop limits C-tier; Hanson predicts institutional absorption not local FOOM.

Evidence:

  • hanson variant human action — 60/40 default mix; misalign tail 25%→10%
  • Anthropic RSI §3 — Amdahl’s law internal
  • Node 12 Hanson table — L1 B-plateau 45% default / 55% Hanson

Analogue: Hanson–Yud 2008 — speed after AGI disputed, not whether agents improve.

Would update if: Single lab >50% frontier FLOPs + internal RSI >10× without sharing — Hanson falsified.

Conf: M


16. Hanson variant — L3/L4 first increases Hanson weight

Claim: If L3 embodied or L4 cyber/finance fires first, P(Hanson mix >50%) = 0.55 — slow takeoff, liability/insurance absorption.

Why: Physical/finance tails = whimper/severe before extinction; institutions already regulate OT/finance better than alignment.

Evidence:

  • Node 12 Hanson integration table — L3 40/60 default/Hanson; L4 50/50
  • Hanson variant node table — N8 modal = Hanson absorption
  • Node 12 non-extinction reweight — L3 +3–5pp whimper

Analogue: Industrial robot safety standards — decades of OSHA absorption.

Would update if: Embodied incident triggers federal training moratorium — anti-Hanson.

Conf: M


Claim: L2 bio tail uses BMIA/synthesis venue — orthogonal to Node 7 compute KYC/BIS — so bio can fire first while cloud FLOPs unchanged.

Why: GeneBreaker, Tier 2 emerging, screening politics (Node 2 P(BMIA)=0.45) — different Kingdon stream than agent labor (Node 1).

Evidence:

  • node2 — BMIA, Trigger E, Tier 2→3
  • Node 12 cross-node matrix — “Tier 3 after cloud C” vs “Front-run screening failure”
  • Node 10 — hidden near-miss delays BMIA → L2 first-tail risk

Analogue: Fentanyl precursors — chemistry chokepoint ≠ internet chokepoint.

Would update if: BMIA passes before Tier 3 + cloud still B-tier → L2 , extinction 2pp (Node 12 falsifier).

Conf: M


Claim: L1 cloud RSI co-locates with Node 7 triopoly (AWS/Azure/GCP) — same hyperscaler chokepoint for FLOPs and agent R&D loops; Node 7 Node 12 but correlates ρ≈0.5+ on L1 modal.

Why: Stargate, Korea 260k GPU, Genesis — total L1 capacity; export controls slow not stop B-tier. Embodied not on triopoly.

Evidence:

Analogue: Cloud concentrates digital RSI; factories distribute embodied.

Would update if: Major lab trains >50% cycles off hyperscaler (sovereign-only) → L1/N7 decouple.

Conf: M


Claim: Node 8 fires capability 2025 H2–27; policy 2027–29 — can front-run Node 4 C10 without changing Ci cloud spine.

Why: Monterrey OT, Mythos agent fraud, Interpol stats — L4 first-tail 0.10; T8-B vetting tail overlaps Node 1 GUARD.

Evidence:

Analogue: Ransomware wave — cyber policy before AGI governance mature.

Would update if: Node 8 Trigger E before 2027-H1 with >2 mo media → L4 .

Conf: M


20. Branch reweight — L1 cloud resolves (+2–4pp extinction)

Claim: Resolving L1 as first tail: Δw_m −5pp, Δw_b −3pp, Δw_x +8pp → P(extinction) ≈21–22% (from ~17% doom baseline (emergent sim)) — apply one column only.

Why: Cloud RSI first misalignment branch (25%×40%) if no halt; worked example in Node 12 parent.

Evidence:

  • Node 12 reweight formulas + worked example
  • (internal note) — branch model
  • Crosscut §6 — “+3–5pp misalign if cloud 10+ pp faster than bio”

Analogue: Bayesian model selection — one locus resolution, not stack.

Would update if: Full Phase 4b re-stitch with correlation discount.

Conf: M


21. Branch reweight — L3 embodied resolves (−2–3pp extinction)

Claim: L3 first tail: Δw_m +5pp, Δw_x −5pp → P(extinction) ≈15.5% short-run extinction, whimper/severe.

Why: Hanson-compatible; kinetic harm legible but not paperclip path; factory vetting not training pause.

Evidence:

  • Node 12 worked example L3
  • Hanson variant — embodied whimper to 28%
  • Node 1 physical wave — GUARD 12–14% embodied tail

Analogue: Industrial automation — harm without extinction.

Would update if: Embodied RSI compounds to C-tier on factory floor before cloud C — rare joint tail.

Conf: M


22. AI 2027 mapping — L1-only RSI vs our weights

Claim: AI 2027 scenario is L1-only for RSI compounding (Agent-2→4, whistleblower Oct 2027); our 22% bio-first + 12% embodied-first are deliberate deviations — physical tails can front-run narrative order.

Why: Scenario useful for L1 timing; insufficient for sim path reweighting. Tracker Behind on Agent-3 superhuman coder (~0.70×).

Evidence:

Analogue: IEA scenarios — one demand narrative, multiple supply paths.

Would update if: AI 2027 adds explicit bio-first beat with dates → reconcile weights.

Conf: M


23. Physical limits — cosmic scale does not bind Node 12 horizon

Claim: Even L1 C-tier cloud RSI does not break Landauer/Bekenstein on 2030 policy horizon — Tier 1 solar-system optimizer still 10¹–10³ yr engineering (superintelligence physical limits §2).

Why: Node 12 is Earth-surface first-tail selector, not galaxy-fill timing. Energy/data caps slow C-tier, not block B-tier on 2026–28 window.

Evidence:

Analogue: Moore’s Law eventually hits physics — after policy-relevant RSI tiers.

Would update if: Datacenter build hard-stopped by grid/energy policy 2027 → L1 , L3/L4 .

Conf: H (for 2030 scope)


24. Conditional Ci — capability stall shifts locus weights

Claim: If global Ci ≤ C6 by 2030, L1 cloud first-tail to 0.35; L2 bio to 0.32; L4 cyber to 0.14 — cloud RSI narrative overpriced.

Why: Without cloud C-tier, physical-layer tails gain share; Hanson mixture . Tracker Behind/Emerging on several C7+ beats.

Evidence:

  • Node 12 conditional Ci table
  • (internal note) — mixed Ahead/Behind
  • Node 12 falsifier — METR flat <8h through 2028

Analogue: Clean energy — if solar stalls, coal reweights in mix scenarios.

Would update if: C9+ Confirmed on tracker 2027 → restore L1 0.52+.

Conf: M


25. L5 multipolar — Yud-coordination-fail sub-variant

Claim: L5 resolution → misalign branch +17pp weight delta; P(extinction) 22–24%; 85% default / 15% Hanson mixture — anti-Hanson.

Why: Simultaneous salience without ordering = coordination failure — Yudkowsky tail 30–40% sub-variant in crosscut §6.

Evidence:

  • Node 12 TAIL-L5 branch — +4–6pp total
  • Hanson falsifiers — kinetic datacenter strike
  • Node 3 + Node 9 — multipolar diplomacy fails

Analogue: 1914 July crisis — linked events, no single first cause in policy.

Would update if: Clear ordering established post-hoc → collapse L5 into L1–L4.

Conf: L


26. Open-weight (Node 5) — accelerates L1 + L2 tamper

Claim: Node 5 open-weight tails accelerate both L1 cloud tamper and L2 bio cascade — does not change first-tail partition unless T5-B bio near-miss public before cloud C.

Why: Persona vectors, GeneBreaker on open weights — dual locus acceleration; Node 12 cross-node matrix.

Evidence:

  • node5 — T5-B bio tamper
  • Crosscut §6 Node 5 link
  • code/typebits/ — persona-axis steering (Node 4 ext overlap)

Analogue: Dual-use biotech export — two pathways, one material.

Would update if: Open-weight bio FM release precedes next cloud frontier → L2 .

Conf: M


27. p(doom) summary — branch selector discipline

Claim: Node 12 point estimates: L1 21%, L2 20%, L3 16%, L4 19%, L5 23% vs ~17% doom baseline (emergent sim) — do not add linearly to N7–N11 node deltas; apply one reweight when locus evidence resolves.

Why: Methodology guardrail from Node 12 parent + my pdoom scenario branches.

Evidence:

  • Node 12 p(doom) summary table
  • correlation matrix — avoid multiplying marginals
  • Crosscut §9 rank #3 RSI locus ±3–6pp

Analogue: Scenario planning — pick one world, not sum.

Would update if: User completes Phase 4b full re-stitch with mixture α from Hanson table.

Conf: M



Update log

DateChange
2026-07-04Initial — 27 evidenced sections from crosscut §6 + Hanson + N2/7/8 links