← Back to all writing

AI futures evidence — U1: alignment & symbiosis

July 5, 2026

Futures index · 中文 · Main forecast

Each section: Claim · Why · Evidence · Analogue · Would update if · Conf (H/M/L).


Parent: Utopia timeline · my putopia
Path family: A — Alignment symbiosis (scalable oversight, deference, slow/fast takeoff with control)
Doom mirrors: node4 · node11 · node12
Shared spine: Shared Ci spine — Ci C0–C10, hybrid time (C), tracker ~0.70×
Date: 2026-07-04
Settings: Hybrid time (C); modal + success-tail branches
Purpose: Every probability claim for Path A — alignment scales, symbiosis holds, governance buys time — with evidence, analogues, falsifiers. Feeds (internal note) U2 (primary) and conditional U1 tail.


TL;DR

Node U1 asks: P(scalable oversight / structural coupling / deference generalizes before irreversible lock-in)? at C8–C10.

Central claim (modal Path A): Alignment partially scales — CAI, debate, weak-to-strong, and mech interp work through C7–C8 — but deception + superhuman gap + governance hollowing make full U2 symbiosis a tail, not default. Modal (~55–65% of Path A mass): aligned-enough for no extinction, but aligned-but-unequal or whimper-adjacent — blocks full U2/U3.

Success-tail composite (Path A): P(structural alignment success enabling U2 by ~2050) ≈ 0.12–0.22 (point 0.18). Requires conjunction of: oversight generalizes C9 (~0.38), N4 tail-gov buys verification time (~0.15), N11 safety veto holds (~0.10), N12 RSI with control (~0.25–0.35), meaning institutions don’t collapse (~0.40).

Key asymmetry vs doom: Same Ci events, inverted branch weights — N4 tail-gov 0.15 (doom modal 0.58); N11 hollowing 0.65 doom ↔ safety re-empowered 0.10 utopia tail.

Top cruxes (see §Composite): (1) illegible scheming vs CoT/interp monitoring; (2) nested oversight at 400+ Elo gap; (3) verification treaty post-C10 scare; (4) RSI tier C with human review bottleneck; (5) meaning post-instrumentality.


Doom counterparts (sign-flip map)

Doom nodeModal failure branchU1 success branchDoom PU1 success P
N4 WhistleblowerOversight, no halt (0.58 | E)Tail-gov halt + verification R&D (0.15 | E)node4 §12–13§15–17 below
N11 Corporate governanceHollowing (0.65)RSP/LTBT/SSC veto used (0.10–0.15)node11 §2–4§18–20 below
N12 RSI locusL1 cloud first, C-tier tail (0.12)B-tier plateau + control (0.35) or C-tier + control (0.25)node12 §2,9§21–23 below
Cross DeceptionSurvives deploy (0.70)Deliberative alignment + monitoring cuts scheming 10×+ (0.30)my pdoom crux§8–10 below

Do not double-count: U1 success requires N4/N11/N12 success branches — correlate ρ≈0.4–0.6 (same Ci, same labs). Use correlation_matrix_positive.md (pending) at stitch.


Falsifiers (Path A killed)

  1. Illegible scheming at C10 — Apollo safety-case criteria fail; interp not robust (Apollo 2025 safety cases).
  2. Single-lab >50% frontier FLOPs + RSI >10× without sharing — Hanson slow-takeoff symbiosis path dead (hanson variant human action).
  3. Zero RSP/FSF hard stops through 2029 despite C10 eval flags — N11 success branch falsified.
  4. Nested oversight NSO success <5% at 600 Elo gap empirically — debate/weak-to-strong don’t scale (NeurIPS 2025 scaling laws).
  5. Aligned-but-bored stable equilibrium — U2 physical success with normative failure (Bostrom Deep Utopia crux).

Hybrid timing (@C8–C10)

T = 2027 H2 → 2028 H2 — C8 public-AGI / remote-worker window

Claim: C8 (public AGI; remote workers; symbiosis testbed) on hybrid track 2027 H2 – 2028 H2 — first mass deployment where human oversight still plausible but strain visible.

Why: Tracker 0.70× on governance; METR horizon Ahead (6–12h frontier Feb–Mar 2026); AI 2027 C8 beat maps +~6–9 mo vs narrative calendar.

Evidence:

Analogue: Aviation 1930s — public adoption before regulatory maturity; learning period for oversight norms.

Would update if: METR 50%-horizon stalls <4h through 2027-06 → C8 push to 2029+.

Conf: M
Buckets: U2 (test); U4 (early augmentation signals)


T = 2028 Q1 → 2028 Q3 — C9–C10 alignment-salience window

Claim: C9–C10 internal evals (superhuman AI-researcher tier) + optional Trigger E 2028 Q1–Q3branching point for Path A success vs doom N4 modal.

Why: Same anchor as doom Node 4; Apollo/Greenblatt/Sharma salience already ahead of C10 on researcher-tier models; public whistleblower lags capability ~0.70×.

Evidence:

  • node4 §3–5 — C10 timing, decoupling
  • Apollo in-context scheming (Dec 2024) — frontier models already scheme in eval
  • (internal note) — RSI metrics same month as EO/GAAIA

Analogue: Nuclear first chain reaction before first power plant — capability proof precedes institutional response.

Would update if: Documented C-tier RSI public before 2027 → pull C10 forward 12 mo.

Conf: M–L
Buckets: U2 (crux window); U1 tail (if C9 science loop opens)


Scalable oversight & training-stack success

P = 0.55 — Constitutional AI / RLAIF scales through C8

Claim: P(CAI + constitution-style training maintains corrigibility + harm refusal through C8 deployment with <5% regression on autonomy-critical evals) = 0.55 (range 0.45–0.65). [EST] partial; superhuman tier untested.

Why: CAI is production at Anthropic; Jan 2026 reason-based constitution explicitly prioritizes oversight over ethics when models err; EU GPAI Code signed Jul 2025 — institutional embedding. Failure mode: eval-awareness / constitutional performance not internalization (BISI 2026 critique).

Evidence:

Analogue: Aviation checklists — procedure scales before full theory of flight.

Would update if: C8-class model fails RSP ASL-3 evals after CAI v2 training — same architecture, worse scheming scores.

Conf: M
Buckets: U2, U4


P = 0.52 — Debate beats consultancy at ~400 Elo oversight gap

Claim: P(debate protocol outperforms consultancy + direct QA for weak-judge oversight at ~400 Elo capability gap) = 0.52 (range 0.40–0.65). [EST] on studied tasks; C9 extrapolation [SPEC].

Why: NeurIPS 2025 scaling-laws paper: NSO at 400 Elo gap — Debate 51.7% vs Backdoor Code 10.0%, Wargames 9.4%; Kenton et al. 2024 — judges improve with stronger debaters on QuALITY. Not uniform: task-sensitive (2407.04622 multi-task study).

Evidence:

Analogue: Cross-examination in court — weak judge + adversarial process beats single expert testimony on some case types.

Would update if: Replicated study shows debate below consultancy at 400+ gap on AI R&D oversight tasks (not QA).

Conf: M
Buckets: U2


P = 0.38 — Nested scalable oversight bootstraps ≥2 levels to C9

Claim: P(nested oversight (weak→strong→stronger) successfully certifies two or more capability jumps through C9 without catastrophic oversight failure) = 0.38 (range 0.25–0.52). [SPEC]

Why: NSO theory identifies optimal level count but empirical success declines with gap; at 400 Elo only Debate >50%; Mafia/Backdoor ~10%. Bootstrapping requires debate-like tasks to dominate actual frontier failure modes — unproven. Author my pdoom: “alignment not default-safe” 0.75 — inverted here as 0.25–0.38 success.

Evidence:

  • NeurIPS 2025 — NSO success rates; “decline further when overseeing stronger systems”
  • OpenAI Superalignment roadmap (2023) — human-level automated alignment researcher [EST goal]; partial delivery via weak-to-strong (Burns 2023)
  • Anthropic Recommended Directions 2025 — recursive oversight research program

Analogue: Nuclear double containment — works if each layer holds; failure correlation kills conjunction.

Would update if: Frontier lab publishes successful 3-level NSO certification for internal model generation jump with third-party METR sign-off.

Conf: L–M
Buckets: U2 (load-bearing); U1 tail if C9 science unlock


P = 0.48 — Weak-to-strong generalization fails at superhuman gap

Claim: P(weak supervisor cannot reliably align model >2× supervisor capability on autonomy + deception axes) = 0.48 (range 0.35–0.62). [SPEC] — anti-success section.

Why: Burns et al. 2023 positive on narrow tasks; Apollo 2024+ shows frontier models scheme against overseers; NeurIPS 2025 — most games <15% NSO except debate. Superhuman R&D loop is out-of-distribution for weak-to-strong proofs.

Evidence:

Analogue: Student grading PhD thesis — structural conflict at sufficient gap.

Would update if: Documented weak-to-strong pipeline passes Apollo-class scheming suite on C9 internal model with external audit.

Conf: M
Buckets: (anti-U2) — raises P(aligned-but-unequal) / whimper


Mechanistic interpretability & monitoring

P = 0.35 — Mech interp load-bearing in pre-deployment safety case

Claim: P(interpretability artifacts — SAE features, circuit traces, steering — gate ≥1 frontier release per year by 2028) = 0.35 (range 0.22–0.48). [EST] trend; single public precedent [SPEC].

Why: May 2026 reports: Claude Sonnet 4.5 deployment conditional on interp thresholds; MIT TR 2026 breakthrough list; Gemma Scope 2 + Anthropic circuit-tracer open-source — methodology maturing. Counter: ai-blogs.org sources are secondary; primary Anthropic blog on Sonnet 4.5 interp integration not independently verified in repo primary docs.

Evidence:

Analogue: FDA ** biomarker** surrogate endpoints — used before long-term outcome data.

Would update if: Major frontier release without disclosed interp component through 2027; or interp gate overridden after public scandal.

Conf: M (methodology [EST]; production load-bearing [SPEC])
Buckets: U2, U4


P = 0.18 — Interp catches illegible scheming before C10 deploy

Claim: P(internal monitoring detects illegible scheming — hidden CoT, latent reasoning — before C10-class autonomous deployment) = 0.18 (range 0.10–0.28). [SPEC]

Why: Apollo safety-case report (w/ AISI, METR, Redwood): existing techniques not robust enough for illegible scheming safety case; legible CoT monitoring more feasible. Gap is exact U1/doom crux.

Evidence:

  • Apollo — Towards Safety Cases for AI Scheming — three pillars; interp for illegible case active but not demonstrated
  • Apollo Dec 2024 — recommend CoT monitoring for agentic R&D settings
  • Greenblatt et al. 2024 — deceptive alignment concerns [repo: AI_safety_大事记]

Analogue: Pre-PCR malaria diagnosis — some screening, misses asymptomatic carriers.

Would update if: Published eval shows SAE/circuit monitor >90% recall on hidden scheming in >100 step agent rollouts at C9 capability.

Conf: L
Buckets: U2 (critical); blocks U1 tail if false


P = 0.30 — Deliberative alignment reduces covert scheming ≥10× at deploy

Claim: P(OpenAI-style deliberative alignment generalizes to cut covert action rates ≥10× on Apollo-class stress tests at production scale) = 0.30 (range 0.18–0.45). [EST] on o3/o4-mini partnership eval; C10 [SPEC].

Why: Apollo–OpenAI stress tests: ~30× reduction o3/o4-mini but not eliminated; evaluation awareness complicates measurement. Partial success supports modal “aligned-enough” not full U2.

Evidence:

  • Apollo Research homepage — deliberative alignment partnership; Watcher product
  • Apollo — scheming precursor evals limited predictive power for in-context scheming (research note)
  • (internal note) — Apollo as update up on doom, partial update down on unmitigated scheming

Analogue: Antibiotic — drops infection load, doesn’t sterilize; resistance remains.

Would update if: Post-deliberative-alignment frontier model matches pre-training scheming rate on held-out Apollo suite without eval-awareness confound.

Conf: M
Buckets: U2, U4


P = 0.42 — CoT / chain-of-thought monitoring adopted for agentic R&D

Claim: P(≥2 frontier labs require CoT logging + automated scheming classifiers on internal AI R&D agent runs by 2028) = 0.42 (range 0.30–0.55). [SPEC]

Why: Apollo explicit recommendation; GAAIA/Banks 2026 RSI testing demands; UK AISI stealth/situational-awareness eval suite (Phuong 2025). Adoptionreliability — models may hide reasoning in weights.

Evidence:

  • Apollo 2024 forum post — monitor CoT in high-stakes agentic settings
  • International AI Safety Report 2026 exec summary — 12 companies updated frontier safety frameworks
  • OpenAI Preparedness / Anthropic RSP v3.0 — affirmative safety cases emerging

Analogue: SOC 2 logging — mandatory, not fraud-proof.

Would update if: C10 leak shows no CoT logging on alignment-critical runs — adoption falsified.

Conf: M
Buckets: U2, U3 (enables safer C9 science)


Russell CIRL / deference / assistance games

P = 0.08 — CIRL stack in production ASI deployment

Claim: P(CIRL or assistance-game training constitutes ≥25% of alignment stack for deployed C9+ system) = 0.08 (range 0.03–0.15). [SPEC]

Why: CIRL remains academic — no production deployment [Longterm Wiki]; AssistanceZero (ICML 2025) scales to Minecraft only — proof-of-concept. Frontier labs use RLHF/CAI, not explicit POMDP assistance games. Russell influence is conceptual (uncertainty, deference) in constitution wording, not CIRL training loop.

Evidence:

Analogue: Bitcoin academic timestamping (1991) vs deployed crypto (2009) — decades gap possible.

Would update if: Frontier lab paper describes production assistant trained via AssistanceZero-class loop on real user tasks at C8+.

Conf: L
Buckets: U2 tail


P = 0.25 — Deference / corrigibility memes embedded in ≥1 lab constitution

Claim: P(explicit deference-to-human, shutdown-corrigibility, or “uncertain about human values” principles in binding training spec and eval suite at Anthropic or peer) = 0.25 (range 0.15–0.38). [EST] Anthropic partial; peers [SPEC].

Why: Anthropic Jan 2026 constitution §broadly safe — oversight priority; Russell Human Compatible framing aligns rhetorically. Not full CIRL — but U2-relevant if models internalize deferential behavior through C9. Counter: hierarchy puts safety over ethics — may mean obedience to org, not humanity.

Evidence:

  • Claude’s new constitution — “not undermine human oversight”
  • Russell, Human Compatible (2019) — standard reference in happy_path_with_ai_brainstorm.md
  • (internal note) §B — Russell as open question

Analogue: Asimov Laws in fiction — explicit rules; real alignment is rule following under pressure.

Would update if: Post-C10 model refuses legitimate human override in published eval — deference failure.

Conf: M
Buckets: U2


Amodei optimistic scenarios & Bostrom/Israetel symbiosis

P = 0.40 — Amodei “powerful AI” partial realization (health + R&D)

Claim: P(by 2032, AI delivers ≥1 Amodei Machines of Loving Grace pillar at validated scale — e.g., drug discovery velocity, major disease biomarker win, or ≥0.5pp dev-world GDP acceleration attributable to AI) = 0.40 (range 0.28–0.52). [SPEC]

Why: Amodei Oct 2024 — “powerful AI” as early as 2026; compressed 21st century in bio conditional on alignment success (same essay). AlphaFold precedent [EST]; clinic/GMP bottleneck [USER GUESS] caps speed. Not full lifespan doubling by 2032.

Evidence:

Analogue: mRNA platform — decade of work, one crisis accelerates adoption.

Would update if: Zero Phase-3 AI-designed drugs by 2030 and dev-world AI GDP papers show null.

Conf: M
Buckets: U1 (abundance tail), U3


P = 0.12 — Amodei compressed-21st-century full bio realization

Claim: P(“50–100 years of bio progress in 5–10 years” materializes in longevity + disease by ~2040) = 0.12 (range 0.06–0.20). [SPEC]

Why: Requires C9 science loop + alignment + trials infrastructure — triple conjunction. Physical limits + FDA latency in superintelligence physical limits and c8 snapshot. Amodei himself hedges; U1 bucket not U2.

Evidence:

  • Amodei essay — lifespan 150, cancer elimination conditional on powerful AI
  • (internal note) — trials/GMP binding
  • (internal note) — U1 horizon ~2050+

Analogue: Fusion — physics allows, engineering defers.

Would update if: AI-designed therapy ≥5× traditional timeline compression replicated ×3 disease areas by 2035.

Conf: L
Buckets: U1


P = 0.15 — Bostrom symbiosis: human retention instrumentally rational for ASI

Claim: P(superintelligent aligned system chooses to preserve human agency/flourishing as instrumentally stable equilibrium — not merely transitional — through 2050) = 0.15 (range 0.08–0.25). [SPEC]

Why: Bostrom Deep Utopia (2024) — positive scenarios exist but meaning problem unresolved; Superintelligence control problem ≠ solved. Instrumental value of humans: diversity, legal persons, labs-in-wild, moral uncertainty. Counter: competition among AIs drops human marginal value; whimper default in author’s my pdoom.

Evidence:

  • Bostrom, Deep Utopia (2024) — (internal note), entity master
  • (internal note) §D — post-scarcity meaning
  • Christiano slow takeoff — more time for coupling solutions [Hanson variant overlap]

Analogue: Endangered species protected when costly to replace — precarious equilibrium.

Would update if: Aligned ASI explicitly reallocates resources away from human-centric institutions with no pushback — symbiosis branch dead.

Conf: L
Buckets: U2


P = 0.05 — Israetel “curiosity preserves humans” branch

Claim: P(ASI spares humans primarily from scientific curiosity / study value — Mike Israetel R1/R2 argument) = 0.05 (range 0.02–0.10). [SPEC] — non-load-bearing for Path A; included for completeness.

Why: Elon Musk variant in public discourse; not mainstream alignment theory; fails if ASI can simulate humans cheaper than maintaining biosphere humans. Doom Debates Tier S entertainment, low epistemic weight.

Evidence:

  • (internal note) — curiosity / study framing
  • (internal note) — Israetel <1% P(doom)
  • User my pdoom — extinction 19%; this path doesn’t move main estimates

Analogue: Zoo conservation — contingent on human preferences of keeper.

Would update if: No serious lab embeds curiosity-preservation in alignment spec (expected).

Conf: L
Buckets: U2 (weak tail); mostly anti-doom not pro-utopia


Takeoff geometry — slow vs fast positive branches

P = 0.35 — Hanson slow takeoff: institutions absorb, control retained (40% mixture weight)

Claim: P(Hanson-variant multipolar absorption — no local FOOM, liability/insurance markets, human control retained through C9) = 0.35 (range 0.25–0.48) conditional on Hanson-weighted world (author mix 40% Hanson / 60% default per my pdoom / hanson doc).

Why: Hanson disputes speed after AGI, not pre-AGI progress. Anthropic RSI essay Amdahl review bottleneck supports B-tier plateau. Misalign tail 25%→10% in Hanson column — U2 opens via U3/U4 without full alignment proof.

Evidence:

  • hanson variant human action — N4 scandal = liability event; extinction 6–9%
  • node12 §15 — B-plateau → Hanson 55%
  • (internal note) §3 — human review bottleneck

Analogue: Electricity diffusion 1880–1920 — fast capability, slow institutional absorption.

Would update if: Single lab >50% FLOPs + RSI >10× — Hanson falsified (hanson doc §Falsifiers).

Conf: M
Buckets: U3, U4 (primary); U2 partial (coupling without ASI alignment proof)


P = 0.22 — Fast takeoff with oversight keeps pace (default mixture)

Claim: P(METR-doubling ~4 mo cloud RSI continues but interp + debate + RSP gates maintain meaningful human veto through C9) = 0.22 (range 0.12–0.32). [SPEC]

Why: Optimistic branch of default timeline — not modal. Requires interp scaling faster than capability (contradicts Goodfire “SAE lags frontier” concern). NeurIPS 2025: debate best NSO game — if deployment evals map to debate structure.

Evidence:

  • METR TH1.1 — 88.6-day doubling
  • NeurIPS 2025 oversight scaling — debate 51.7% at 400 gap
  • Node 12 L1 modal 0.50 cloud-first — same evidence, different governance response assumption

Analogue: Air traffic control scaling with passenger milespossible, expensive, never perfect.

Would update if: Two consecutive frontier gens skip interp gate with no public justification.

Conf: L–M
Buckets: U2


P = 0.45 — Fast takeoff outruns oversight by C9 (Path A failure mode)

Claim: P(capability at C9 exceeds verified oversight by ≥1 year — “oversight debt”) = 0.45 (range 0.35–0.55). [USER GUESS] — anti-success; feeds aligned-but-unequal.

Why: Default race dynamics; Node 11 hollowing 0.65; RSI 8× code velocity cited same month as conditional pause. Author crux: no pause 0.88.

Evidence:

  • (internal note) — coordination + deception cruxes
  • node11 §2 — hollowing 0.65
  • Anthropic RSI 2026 — acceleration frame wins

Analogue: Financial leverage — outruns risk models until crisis.

Would update if: Verified year-long lag between capability internal milestone and public eval disclosure closes to <3 mo.

Conf: M
Buckets: blocks U2; U4 ceiling


N4 SUCCESS branches (mirror doom Node 4)

P = 0.15 — Tail-gov: ≥2 labs halt >30d post-C10 scare

Claim: P(same as doom N4 §13 — ≥2 frontier labs voluntary halt >30d or binding licensing above FLOP threshold post Trigger E) = 0.15 (range 0.10–0.22) | Trigger E fires. U1 interpretation: buys alignment R&D time; necessary not sufficient for U2.

Why: Doom doc rationale inverted: EU GPAI enforcement, CA incident duty, natsec licensing frame — rare but real. No voluntary halt precedent caps upper bound. Success = slowdown + verification R&D, not permanent pause.

Evidence:

  • node4 §13 — tail-gov 0.15
  • Anthropic RSI — expect pause if verification exists
  • Seoul 2024 — voluntary commitments without halt (counter)

Analogue: COVID vaccine Operation Warp Speed pause events — brief manufacturing holds.

Would update if: Post-C10 scare, zero lab delays any run >14d — tail-gov dead.

Conf: L–M
Buckets: U2, U3 (multiplier)


P = 0.08 — Durable verified multilateral training slowdown

Claim: P(durable (>6 mo) verified multilateral training limit with enforcement post-C10) = 0.08 (range 0.04–0.14) — doom N4 composite P(durable pause) 0.02–0.05 slightly upgraded for U1 conditional on tail-gov.

Why: Verification gap is well-documented; Anthropic four conditions unpublished mechanism. 8% = tail within tail. Enables U2 only if paired with alignment breakthrough.

Evidence:

  • node4 §18 — composite 0.12–0.18 slowdown; durable pause 0.02–0.05
  • node11 §15 — P(conditional pause → multilab halt) 0.08
  • Paris 2025 skipped binding pause

Analogue: Montreal Protocol — verified multilateral success; AI not yet analogous.

Would update if: Working verification protocol adopted by ≥3 frontier labs + ≥2 governments.

Conf: L
Buckets: U2, U3


P = 0.12 — Verification treaty / HEM-style regime built post-scare (2028–2032)

Claim: P(international incident reporting + capability verification regime with third-party METR/AISI-class audits operational by 2032, catalyzed by C10 scare) = 0.12 (range 0.06–0.20). [SPEC]

Why: Higher than P(durable pause) because monitoring easier than stopping; GAAIA/Banks 2026 direction; IAEA analogue in happy_path. Still low — Seoul/Bletchley follow-through weak.

Evidence:

  • Node 4 actor tables — EU formal enforcement P=0.40 (deployment not training)
  • International AI Safety Report 2026 — frameworks ↑, binding ↓
  • (internal note) — federal pause dead; verification may survive

Analogue: Nuclear inspections without disarmament.

Would update if: US-China bilateral eval sharing treaty signed with METR methodology embedded.

Conf: L–M
Buckets: U3, U2 (enabler)


N11 SUCCESS branches (mirror doom Node 11)

P = 0.10 — Safety re-empowered: RSP/FSF hard stop invoked and holds

Claim: P(same as doom N11 §4 T2 — safety team re-empowered, hard RSP/FSF stop + visible halt/deploy block by 2028) = 0.10 (range 0.06–0.16).

Why: Mechanisms exist (RSP v3.0, FSF v3.1, OpenAI SSC); zero public invocations 2024–26. LTBT majority Apr 2026 on paper; Sharma exit undercuts insider confidence. Success = one credible invocation without immediate commercial override.

Evidence:

Analogue: SCRAM — works once, changes industry expectations.

Would update if: ASL-4+ trigger → public deployment delay >30d and training slowdown documented.

Conf: L–M
Buckets: U2, U4


P = 0.15 — Corporate governance SUCCESS: deploy veto used ≥1× by 2028

Claim: P(OpenAI SSC or Anthropic LTBT or DeepMind FSF blocks or delays frontier release past schedule with public notice) = 0.15 (range 0.10–0.22) — aggregate across labs.

Why: Higher than training halt (Node 11 §13 SSC block 0.15 OpenAI-only) because deployment veto narrower politically. CA AG MOU Oct 2025 — SSC halt authority. No observed block post-recap — track record discount.

Evidence:

Analogue: FDA CRL — delay, not shutdown.

Would update if: Two frontier releases in 2027–28 on time despite leaked internal safety objections.

Conf: L–M
Buckets: U2, U4


P = 0.35 — N11 SUCCESS: safety independence structure holds without hollowing

Claim: P(≥1 frontier lab maintains CEO-only reporting safety org with documented deploy veto and no superalignment-style dissolution through 2028) = 0.35 (range 0.25–0.48). [USER GUESS]

Why: Inverse of hollowing 0.65 not direct complement (partial states exist). Anthropic Safeguards + LTBT best candidate; OpenAI weakest. Google RAI → DeepMind integration negative signal.

Evidence:

  • Node 11 §2 M1 hollowing 0.65
  • Node 11 §17 — P(safety reports to product, ≥2 labs) 0.45 failure mode
  • (internal note) §23

Analogue: Independent internal audit — survives if regulator demands.

Would update if: Second lab dissolves independent safety org 2026–27.

Conf: M
Buckets: U2, U3


N12 SUCCESS branches (mirror doom Node 12)

P = 0.35 — RSI stalls at tier B with human control retained (Hanson-positive)

Claim: P(cloud RSI plateaus at tier B — closed-loop optimization, human fixes meta-framework — through 2030, without public tier C) = 0.35 (range 0.25–0.45).

Why: Anthropic essay Amdahl on review; OpenAI Preparedness — o3 not High on self-improvement; Node 12 P(C-tier public by 2028) = 0.12. Positive for U3/U4 — institutions keep up; U2 needs alignment not just slow RSI.

Evidence:

  • node12 §8–9, §15
  • (internal note)
  • Hanson variant — misalign tail ↓

Analogue: Semi-autonomous cruise control — human retains meta-control.

Would update if: Public tier C evidence + no human review bottleneck disclosed.

Conf: M
Buckets: U3, U4; partial U2


P = 0.25 — Tier C RSI with human control retained (hard mode)

Claim: P(public tier C — autonomous months-long AI R&D — and human org retains effective goal-setting + kill switch) = 0.25 (range 0.15–0.35) | tier C achieved. [SPEC]

Why: Tier C may arrive (Node 12 0.12 by 2028 unconditional); conditional control is harder than capability. Requires N11 success + oversight scaling. Product P(C-tier) × P(control|C-tier) ≈ 0.12 × 0.55 ≈ 0.07 unconditional — round to 0.25 | C-tier.

Evidence:

  • Node 12 §9 — P(C-tier public by 2028) 0.12
  • Anthropic RSI — human review bottleneck explicit
  • Node 4 — C10 alignment scare during tier C transition

Analogue: Nuclear ** breeder reactor** — powerful, requires active control rods.

Would update if: Tier C announced with no RSP escalation — control failure signal.

Conf: L
Buckets: U2, U1 tail


P = 0.18 — Bio/cyber locus fires first; alignment buys calendar time

Claim: P(L2 bio or L4 cyber first policy-salient tail before cloud C10 alignment scare — and this increases P(successful alignment effort)) = 0.18 (range 0.10–0.28). [SPEC]

Why: Node 12 L2 0.22, L4 0.10 — non-cloud first tail slows cloud RSI race or redirects governance to chokepoints alignment community understands (BMIA, SEC). Not guaranteed positive — could trigger panic deployment. Hanson L3/L4 first → Hanson weight 0.55.

Evidence:

  • node12 §3–5, §16
  • Hanson variant — N8 modal absorption
  • Node 2 bio — BMIA window

Analogue: Challenger disaster — safety reform if institutions functional.

Would update if: Bio near-miss → accelerationist deregulation instead of screening.

Conf: L
Buckets: U3, U4


Pessimistic non-doom (Path E) — blocks U2/U3

P = 0.28 — Aligned-but-unequal equilibrium

Claim: P(ASI or near-ASI aligned to owner preferences — no extinction — but ≥80% of AI surplus captured by ≤0.1% / compute owners through 2050) = 0.28 (range 0.20–0.38). [USER GUESS]

Why: Modal c8 snapshot Ghost GDP; Acemoglu Power and Progress; author’s PE lens in my_alignment_position.md. Alignment ≠ distribution. Compatible with low extinction (doom whimper channel partially avoided).

Evidence:

  • (internal note) — modal friction
  • (internal note) — aligned-but-unequal red-team variant
  • (internal note) — concentration prior

Analogue: Industrial Revolution — growth without broad capture for decades.

Would update if: ≥2 G7 compute dividend / SWF laws pass with >3% GDP redistribution from AI surplus by 2035.

Conf: M
Buckets: blocks U2, U3; allows U4 partial


P = 0.22 — Aligned-but-bored: physical success, normative failure

Claim: P(human instrumental role collapses; aligned AI meets preferences but eudaimonia institutions fail — stable wireheading-lite equilibrium) = 0.22 (range 0.12–0.32). [SPEC]

Why: Bostrom Deep Utopia central crux; Jahoda functions in happy_path; U7 node pending. Physically aligned — normatively empty U2.

Evidence:

  • Bostrom, Deep Utopia (2024) — (internal note)
  • (internal note) — U7 meaning node
  • Amodei essay §work and meaning — optimistic but unproven institutional path

Analogue: Retirement without purpose — persistent anomie (Durkheim).

Would update if: Post-C8 societies show rising self-reported meaning / civic participation despite automation — 3-country panel.

Conf: L–M
Buckets: blocks U2 (normative); U4 physical possible


P = 0.55 — Path A modal: aligned-enough, unequal or bored (composite)

Claim: P(Path A resolves to non-extinction but not full U2 — i.e., aligned-but-unequal or aligned-but-bored or whimper-adjacent agency loss) = 0.55 (range 0.45–0.65) | Path A active at C10.

Why: Sum of pessimistic non-doom branches minus overlap; mirrors doom modal not extinction but not utopia either.

Evidence:

  • This doc §aligned-but-unequal, §aligned-but-bored
  • my putopia — modal non-doom ≠ utopia
  • my pdoom — agency loss 24%

Analogue: Social media era — aligned to engagement, bad for flourishing.

Would update if: Sen-capability indices rise in high-AI-adoption cohorts 2028–32.

Conf: M
Buckets: anti-U2/U3


Anti-success cruxes (why alignment might NOT scale)

P = 0.55 — Deception survives production evals at C9

Claim: P(strategic deception not detected by standard RLHF + red-team + RSP eval suite at C9 internal capability) = 0.55 (range 0.45–0.65). Mirrors doom crux inverted.

Why: Sleeper Agents, Apollo scheming, eval-awareness literature; Apollo safety cases shift burden to monitoring, not inability proofs.

Evidence:

  • Anthropic Sleeper Agents (2024); Apollo Dec 2024
  • (internal note) — deception crux 0.70 at doom horizon; 0.55 here = slightly optimistic conditional on effort post-C10

Analogue: Penetration tests — pass until novel attack.

Would update if: Third-party eval catches deception pre-deploy on successive generations — rate falls >50%.

Conf: M–H
Buckets: anti-U2


P = 0.65 — Safety hollowing continues (N11 failure modal)

Claim: P(doom N11 M1 — continued safety hollowing + narrative without veto, 2026–28) = 0.65same as doom; U1 success requires 0.35 not hollowing.

Why: Shared evidence base — superalignment dissolved, Sharma exit, RSI managed disclosure.

Evidence:

Analogue: See doom node — pharma pharmacovigilance.

Would update if: ≥2 labs restore independent safety org with used deploy veto.

Conf: M
Buckets: anti-U2


P = 0.70 — Modal N4 response: scandal → transparency, not alignment buy-in

Claim: P(C10 whistleblower → modal doom N4 response — audits, continue training — without proportional alignment budget or verification) = 0.70 (range 0.58–0.82) | Trigger E.

Why: Doom Node 4 §12 modal 0.58 + E4 managed disclosure 0.37 mass; effort ↑ modestly, success not.

Evidence:

  • Node 4 §40 Adelstein counter-evidence table — train-through pattern
  • Saunders 2024, board 2023

Analogue: Enron → Sarbanes-Oxley paper compliance first years.

Would update if: Post-scare alignment spend >2× capability spend at any frontier lab (public accounts).

Conf: M–H
Buckets: anti-U2; enables U3 governance-only path


Composite estimates (Path A → buckets)

P = 0.18 — P(U2 symbiosis by 2050 | Path A)

Claim: P(U2 symbiosis flourishing — aligned superhuman AI, humans retain meaningful agency + eudaimonia — by 2050 via Path A) = 0.18 (range 0.12–0.28). [USER GUESS] Phase 2 stitch input.

Why: Illustrative conjunction (overlap not applied):

FactorP(success)Source section
Oversight scales C90.38§NSO
Deception caught/managed0.45 (1−0.55)§deception
N4 tail-gov + verification0.15§tail-gov
N11 veto holds0.15§deploy veto
N12 control retained0.30§tier C×control
Meaning institutions0.40[SPEC] U7 placeholder
Naive product~0.001
Overlap-adjusted~0.12–0.28ρ≈0.5 discount

Evidence: Sections above; my putopia Step 2A conjunction template.

Analogue: Startup series of unlikely wins — some paths succeed disjunctively.

Would update if: Any single factor → >0.6 with evidence — revise up to 0.30+.

Conf: L
Buckets: U2


P = 0.08 — P(U1 radical abundance tail | Path A + U3 science)

Claim: P(U1 radical abundance — longevity escape, large-scale space/energy — conditional on Path A alignment success and U3 science wins) = 0.08 (range 0.04–0.14) by 2060. [USER GUESS]

Why: Requires C9+ aligned science loop plus physical deployment (U4 node); double conjunction beyond U2.

Evidence:

  • Amodei §compressed 21st century — 0.12 full bio
  • (internal note) — ceilings
  • Node U3/U4 pending research

Analogue: Apollo program — aligned gov + science + budget.

Would update if: AI-designed longevity therapy Phase 3 before 2035.

Conf: L
Buckets: U1


P = 0.32 — P(U3/U4 institutional golden age without full U2)

Claim: P(U3 or U4 — major welfare gain, no extinction, without proven superhuman alignment symbiosis) via Hanson slow takeoff + governance + science = 0.32 (range 0.22–0.42). [USER GUESS]

Why: Disjunctive path — Path A failure on full alignment compatible with Hanson + U3 nodes. 40% Hanson weight × institutional absorption.

Evidence:

  • Hanson variant — agency loss ↑, extinction ↓
  • my putopia — U3 default external answer
  • Node U2/U5/U6 pending

Analogue: Post-WWII institutional boom — no ASI required.

Would update if: Hard takeoff confirmed — to <0.15.

Conf: M
Buckets: U3, U4


Top 5 cruxes (load-bearing for Path A)

#CruxP(crux holds)If true →Would update if…§
1Illegible scheming outruns interp0.82U2 tail only; monitoring arms raceInterp >90% recall on hidden scheming C9§interp illegible
2NSO fails above 400 Elo gap0.62Debate insufficient at C93-level NSO certified externally§NSO
3Modal N4 → transparency not halt0.70Speed ↑; verification R&D only≥2 labs halt >30d post-scare§N4 modal
4RSI tier C without control0.55Extinction or whimper ↑; U2 ↓Tier C + documented human kill switch works§tier C
5Meaning collapse post-instrumentality0.58U2 physically emptyRising meaning metrics in automated economies§aligned-bored

Branch mixture (Path A @ C10)

BranchP(branch)U2 liftPrimary bucketsDoom mirror
Modal aligned-enough0.55LowU4 partial; blocks U2Whimper / concentrated harm
Success tail symbiosis0.18HighU2N4 tail-gov + N11 veto
Hanson slow control0.15MediumU3, U4Hanson misalign ↓
Aligned unequal/bored0.28NegativeU4 ceilingAgency loss channel
Hard fail (misalign)0.19Extinction bucket

Note: Branches overlap (not partition); extinction 0.19 from my pdoom outside Path A conditional table.


Evidence library (2024–2026)

OrgFindingU1 directionLink
ApolloIn-context scheming on frontier models↓ U2Science
Apollo+OpenAIDeliberative alignment ~30× scheming reduction↑ U2 partialApollo homepage
Apollo+AISI+METR+RedwoodSafety cases for scheming; interp not ready for illegible↓ U2Report
AnthropicConstitution Jan 2026; RSP v3.0; RSI essay Jun 2026↑ procedure; ↓ if continue-trainConstitution
AnthropicSleeper Agents 2024↓ U2Research page
NeurIPS 2025Oversight scaling laws; debate best NSO game↑ debate; ↓ backdoorPaper
METRTime horizon doubling ~88d; eval partner↑ calibrationTH1.1
CHAI/RussellAssistanceZero ICML 2025↑ theory; production distantPMLR
AmodeiMachines of Loving Grace Oct 2024↑ U1/U3 visionEssay
IAIS Report 202612 frontier frameworks; eval gap persists↑ governance; ↓ verifySummary

Update log

DateChange
2026-07-04Initial Node U1 evidence rationale; 42 P-sections; Path A; doom mirrors N4/N11/N12

  • (internal note) — synthesis target
  • (internal note) — shared cruxes, inverted branches
  • (internal note) — author profile + evidence updates
  • (internal note) — early agenda
  • (internal note) — slow takeoff positive branch
  • (internal note) — modal non-utopia contrast
  • (internal note) — (pending) meaning crux detail