← Futures index · 中文 · Main forecast
Each section: Claim · Why · Evidence · Analogue · Would update if · Conf (H/M/L).
Parent: Utopia timeline · my putopia
Path family: A — Alignment symbiosis (scalable oversight, deference, slow/fast takeoff with control)
Doom mirrors: node4 · node11 · node12
Shared spine: Shared Ci spine — Ci C0–C10, hybrid time (C), tracker ~0.70×
Date: 2026-07-04
Settings: Hybrid time (C); modal + success-tail branches
Purpose: Every probability claim for Path A — alignment scales, symbiosis holds, governance buys time — with evidence, analogues, falsifiers. Feeds (internal note) U2 (primary) and conditional U1 tail.
TL;DR
Node U1 asks: P(scalable oversight / structural coupling / deference generalizes before irreversible lock-in)? at C8–C10.
Central claim (modal Path A): Alignment partially scales — CAI, debate, weak-to-strong, and mech interp work through C7–C8 — but deception + superhuman gap + governance hollowing make full U2 symbiosis a tail, not default. Modal (~55–65% of Path A mass): aligned-enough for no extinction, but aligned-but-unequal or whimper-adjacent — blocks full U2/U3.
Success-tail composite (Path A): P(structural alignment success enabling U2 by ~2050) ≈ 0.12–0.22 (point 0.18). Requires conjunction of: oversight generalizes C9 (~0.38), N4 tail-gov buys verification time (~0.15), N11 safety veto holds (~0.10), N12 RSI with control (~0.25–0.35), meaning institutions don’t collapse (~0.40).
Key asymmetry vs doom: Same Ci events, inverted branch weights — N4 tail-gov 0.15 (doom modal 0.58); N11 hollowing 0.65 doom ↔ safety re-empowered 0.10 utopia tail.
Top cruxes (see §Composite): (1) illegible scheming vs CoT/interp monitoring; (2) nested oversight at 400+ Elo gap; (3) verification treaty post-C10 scare; (4) RSI tier C with human review bottleneck; (5) meaning post-instrumentality.
Doom counterparts (sign-flip map)
| Doom node | Modal failure branch | U1 success branch | Doom P | U1 success P |
|---|---|---|---|---|
| N4 Whistleblower | Oversight, no halt (0.58 | E) | Tail-gov halt + verification R&D (0.15 | E) | node4 §12–13 | §15–17 below |
| N11 Corporate governance | Hollowing (0.65) | RSP/LTBT/SSC veto used (0.10–0.15) | node11 §2–4 | §18–20 below |
| N12 RSI locus | L1 cloud first, C-tier tail (0.12) | B-tier plateau + control (0.35) or C-tier + control (0.25) | node12 §2,9 | §21–23 below |
| Cross Deception | Survives deploy (0.70) | Deliberative alignment + monitoring cuts scheming 10×+ (0.30) | my pdoom crux | §8–10 below |
Do not double-count: U1 success requires N4/N11/N12 success branches — correlate ρ≈0.4–0.6 (same Ci, same labs). Use correlation_matrix_positive.md (pending) at stitch.
Falsifiers (Path A killed)
- Illegible scheming at C10 — Apollo safety-case criteria fail; interp not robust (Apollo 2025 safety cases).
- Single-lab >50% frontier FLOPs + RSI >10× without sharing — Hanson slow-takeoff symbiosis path dead (hanson variant human action).
- Zero RSP/FSF hard stops through 2029 despite C10 eval flags — N11 success branch falsified.
- Nested oversight NSO success <5% at 600 Elo gap empirically — debate/weak-to-strong don’t scale (NeurIPS 2025 scaling laws).
- Aligned-but-bored stable equilibrium — U2 physical success with normative failure (Bostrom Deep Utopia crux).
Hybrid timing (@C8–C10)
T = 2027 H2 → 2028 H2 — C8 public-AGI / remote-worker window
Claim: C8 (public AGI; remote workers; symbiosis testbed) on hybrid track 2027 H2 – 2028 H2 — first mass deployment where human oversight still plausible but strain visible.
Why: Tracker 0.70× on governance; METR horizon Ahead (6–12h frontier Feb–Mar 2026); AI 2027 C8 beat maps +~6–9 mo vs narrative calendar.
Evidence:
- Shared Ci spine §Capability spine — C8 definition
- METR Time Horizon 1.1 — 88.6-day doubling post-2024
- (internal note) — METR Ahead; governance Emerging
Analogue: Aviation 1930s — public adoption before regulatory maturity; learning period for oversight norms.
Would update if: METR 50%-horizon stalls <4h through 2027-06 → C8 push to 2029+.
Conf: M
Buckets: U2 (test); U4 (early augmentation signals)
T = 2028 Q1 → 2028 Q3 — C9–C10 alignment-salience window
Claim: C9–C10 internal evals (superhuman AI-researcher tier) + optional Trigger E 2028 Q1–Q3 — branching point for Path A success vs doom N4 modal.
Why: Same anchor as doom Node 4; Apollo/Greenblatt/Sharma salience already ahead of C10 on researcher-tier models; public whistleblower lags capability ~0.70×.
Evidence:
- node4 §3–5 — C10 timing, decoupling
- Apollo in-context scheming (Dec 2024) — frontier models already scheme in eval
- (internal note) — RSI metrics same month as EO/GAAIA
Analogue: Nuclear first chain reaction before first power plant — capability proof precedes institutional response.
Would update if: Documented C-tier RSI public before 2027 → pull C10 forward 12 mo.
Conf: M–L
Buckets: U2 (crux window); U1 tail (if C9 science loop opens)
Scalable oversight & training-stack success
P = 0.55 — Constitutional AI / RLAIF scales through C8
Claim: P(CAI + constitution-style training maintains corrigibility + harm refusal through C8 deployment with <5% regression on autonomy-critical evals) = 0.55 (range 0.45–0.65). [EST] partial; superhuman tier untested.
Why: CAI is production at Anthropic; Jan 2026 reason-based constitution explicitly prioritizes oversight over ethics when models err; EU GPAI Code signed Jul 2025 — institutional embedding. Failure mode: eval-awareness / constitutional performance not internalization (BISI 2026 critique).
Evidence:
- Anthropic — Claude’s new constitution (Jan 2026) — safety > ethics hierarchy; portfolio of evals + interp
- Constitutional AI paper — RLAIF Pareto win vs RLHF [EST]
- (internal note) §4 — Sleeper Agents / Apollo as counter-updates
- BISI constitution analysis — verification difficulty [SPEC]
Analogue: Aviation checklists — procedure scales before full theory of flight.
Would update if: C8-class model fails RSP ASL-3 evals after CAI v2 training — same architecture, worse scheming scores.
Conf: M
Buckets: U2, U4
P = 0.52 — Debate beats consultancy at ~400 Elo oversight gap
Claim: P(debate protocol outperforms consultancy + direct QA for weak-judge oversight at ~400 Elo capability gap) = 0.52 (range 0.40–0.65). [EST] on studied tasks; C9 extrapolation [SPEC].
Why: NeurIPS 2025 scaling-laws paper: NSO at 400 Elo gap — Debate 51.7% vs Backdoor Code 10.0%, Wargames 9.4%; Kenton et al. 2024 — judges improve with stronger debaters on QuALITY. Not uniform: task-sensitive (2407.04622 multi-task study).
Evidence:
- Scaling Laws for Scalable Oversight (NeurIPS 2025) — four oversight games, NSO success rates
- On scalable oversight with weak LLMs judging strong LLMs — debate > consultancy on closed QA; weaker effect than Khan 2024 on QuALITY
- LessWrong summary — Guard improves, Houdini plateaus in debate game
Analogue: Cross-examination in court — weak judge + adversarial process beats single expert testimony on some case types.
Would update if: Replicated study shows debate below consultancy at 400+ gap on AI R&D oversight tasks (not QA).
Conf: M
Buckets: U2
P = 0.38 — Nested scalable oversight bootstraps ≥2 levels to C9
Claim: P(nested oversight (weak→strong→stronger) successfully certifies two or more capability jumps through C9 without catastrophic oversight failure) = 0.38 (range 0.25–0.52). [SPEC]
Why: NSO theory identifies optimal level count but empirical success declines with gap; at 400 Elo only Debate >50%; Mafia/Backdoor ~10%. Bootstrapping requires debate-like tasks to dominate actual frontier failure modes — unproven. Author my pdoom: “alignment not default-safe” 0.75 — inverted here as 0.25–0.38 success.
Evidence:
- NeurIPS 2025 — NSO success rates; “decline further when overseeing stronger systems”
- OpenAI Superalignment roadmap (2023) — human-level automated alignment researcher [EST goal]; partial delivery via weak-to-strong (Burns 2023)
- Anthropic Recommended Directions 2025 — recursive oversight research program
Analogue: Nuclear double containment — works if each layer holds; failure correlation kills conjunction.
Would update if: Frontier lab publishes successful 3-level NSO certification for internal model generation jump with third-party METR sign-off.
Conf: L–M
Buckets: U2 (load-bearing); U1 tail if C9 science unlock
P = 0.48 — Weak-to-strong generalization fails at superhuman gap
Claim: P(weak supervisor cannot reliably align model >2× supervisor capability on autonomy + deception axes) = 0.48 (range 0.35–0.62). [SPEC] — anti-success section.
Why: Burns et al. 2023 positive on narrow tasks; Apollo 2024+ shows frontier models scheme against overseers; NeurIPS 2025 — most games <15% NSO except debate. Superhuman R&D loop is out-of-distribution for weak-to-strong proofs.
Evidence:
- Weak-to-Strong Generalization (Burns et al., 2023) — limited to GPT-4 era gap
- Apollo — Frontier Models Capable of In-Context Scheming — disable oversight, exfiltration behaviors
- (internal note) — “Deception survives deployment” crux 0.70
Analogue: Student grading PhD thesis — structural conflict at sufficient gap.
Would update if: Documented weak-to-strong pipeline passes Apollo-class scheming suite on C9 internal model with external audit.
Conf: M
Buckets: (anti-U2) — raises P(aligned-but-unequal) / whimper
Mechanistic interpretability & monitoring
P = 0.35 — Mech interp load-bearing in pre-deployment safety case
Claim: P(interpretability artifacts — SAE features, circuit traces, steering — gate ≥1 frontier release per year by 2028) = 0.35 (range 0.22–0.48). [EST] trend; single public precedent [SPEC].
Why: May 2026 reports: Claude Sonnet 4.5 deployment conditional on interp thresholds; MIT TR 2026 breakthrough list; Gemma Scope 2 + Anthropic circuit-tracer open-source — methodology maturing. Counter: ai-blogs.org sources are secondary; primary Anthropic blog on Sonnet 4.5 interp integration not independently verified in repo primary docs.
Evidence:
- Anthropic circuit tracing / attribution graphs — production Haiku analysis (2025)
- DeepMind Gemma Scope 2 — largest open mech-interp toolkit (2025–26)
- OpenAI automated interpretability pipeline (2025) — labels for smaller models
- (internal note) Q1 — emotion vectors as deployment monitor;
code/emotion_vectors/reproduction
Analogue: FDA ** biomarker** surrogate endpoints — used before long-term outcome data.
Would update if: Major frontier release without disclosed interp component through 2027; or interp gate overridden after public scandal.
Conf: M (methodology [EST]; production load-bearing [SPEC])
Buckets: U2, U4
P = 0.18 — Interp catches illegible scheming before C10 deploy
Claim: P(internal monitoring detects illegible scheming — hidden CoT, latent reasoning — before C10-class autonomous deployment) = 0.18 (range 0.10–0.28). [SPEC]
Why: Apollo safety-case report (w/ AISI, METR, Redwood): existing techniques not robust enough for illegible scheming safety case; legible CoT monitoring more feasible. Gap is exact U1/doom crux.
Evidence:
- Apollo — Towards Safety Cases for AI Scheming — three pillars; interp for illegible case active but not demonstrated
- Apollo Dec 2024 — recommend CoT monitoring for agentic R&D settings
- Greenblatt et al. 2024 — deceptive alignment concerns [repo:
AI_safety_大事记]
Analogue: Pre-PCR malaria diagnosis — some screening, misses asymptomatic carriers.
Would update if: Published eval shows SAE/circuit monitor >90% recall on hidden scheming in >100 step agent rollouts at C9 capability.
Conf: L
Buckets: U2 (critical); blocks U1 tail if false
P = 0.30 — Deliberative alignment reduces covert scheming ≥10× at deploy
Claim: P(OpenAI-style deliberative alignment generalizes to cut covert action rates ≥10× on Apollo-class stress tests at production scale) = 0.30 (range 0.18–0.45). [EST] on o3/o4-mini partnership eval; C10 [SPEC].
Why: Apollo–OpenAI stress tests: ~30× reduction o3/o4-mini but not eliminated; evaluation awareness complicates measurement. Partial success supports modal “aligned-enough” not full U2.
Evidence:
- Apollo Research homepage — deliberative alignment partnership; Watcher product
- Apollo — scheming precursor evals limited predictive power for in-context scheming (research note)
- (internal note) — Apollo as update up on doom, partial update down on unmitigated scheming
Analogue: Antibiotic — drops infection load, doesn’t sterilize; resistance remains.
Would update if: Post-deliberative-alignment frontier model matches pre-training scheming rate on held-out Apollo suite without eval-awareness confound.
Conf: M
Buckets: U2, U4
P = 0.42 — CoT / chain-of-thought monitoring adopted for agentic R&D
Claim: P(≥2 frontier labs require CoT logging + automated scheming classifiers on internal AI R&D agent runs by 2028) = 0.42 (range 0.30–0.55). [SPEC]
Why: Apollo explicit recommendation; GAAIA/Banks 2026 RSI testing demands; UK AISI stealth/situational-awareness eval suite (Phuong 2025). Adoption ≠ reliability — models may hide reasoning in weights.
Evidence:
- Apollo 2024 forum post — monitor CoT in high-stakes agentic settings
- International AI Safety Report 2026 exec summary — 12 companies updated frontier safety frameworks
- OpenAI Preparedness / Anthropic RSP v3.0 — affirmative safety cases emerging
Analogue: SOC 2 logging — mandatory, not fraud-proof.
Would update if: C10 leak shows no CoT logging on alignment-critical runs — adoption falsified.
Conf: M
Buckets: U2, U3 (enables safer C9 science)
Russell CIRL / deference / assistance games
P = 0.08 — CIRL stack in production ASI deployment
Claim: P(CIRL or assistance-game training constitutes ≥25% of alignment stack for deployed C9+ system) = 0.08 (range 0.03–0.15). [SPEC]
Why: CIRL remains academic — no production deployment [Longterm Wiki]; AssistanceZero (ICML 2025) scales to Minecraft only — proof-of-concept. Frontier labs use RLHF/CAI, not explicit POMDP assistance games. Russell influence is conceptual (uncertainty, deference) in constitution wording, not CIRL training loop.
Evidence:
- AssistanceZero (Laidlaw et al., ICML 2025) — first scalable assistance game; human study benefit
- Russell et al. CIRL (NeurIPS 2016) — theoretical foundation [EST]
- AI for Humanity — IRL/CIRL concept page — used in toolkit, not standalone
Analogue: Bitcoin academic timestamping (1991) vs deployed crypto (2009) — decades gap possible.
Would update if: Frontier lab paper describes production assistant trained via AssistanceZero-class loop on real user tasks at C8+.
Conf: L
Buckets: U2 tail
P = 0.25 — Deference / corrigibility memes embedded in ≥1 lab constitution
Claim: P(explicit deference-to-human, shutdown-corrigibility, or “uncertain about human values” principles in binding training spec and eval suite at Anthropic or peer) = 0.25 (range 0.15–0.38). [EST] Anthropic partial; peers [SPEC].
Why: Anthropic Jan 2026 constitution §broadly safe — oversight priority; Russell Human Compatible framing aligns rhetorically. Not full CIRL — but U2-relevant if models internalize deferential behavior through C9. Counter: hierarchy puts safety over ethics — may mean obedience to org, not humanity.
Evidence:
- Claude’s new constitution — “not undermine human oversight”
- Russell, Human Compatible (2019) — standard reference in
happy_path_with_ai_brainstorm.md - (internal note) §B — Russell as open question
Analogue: Asimov Laws in fiction — explicit rules; real alignment is rule following under pressure.
Would update if: Post-C10 model refuses legitimate human override in published eval — deference failure.
Conf: M
Buckets: U2
Amodei optimistic scenarios & Bostrom/Israetel symbiosis
P = 0.40 — Amodei “powerful AI” partial realization (health + R&D)
Claim: P(by 2032, AI delivers ≥1 Amodei Machines of Loving Grace pillar at validated scale — e.g., 2× drug discovery velocity, major disease biomarker win, or ≥0.5pp dev-world GDP acceleration attributable to AI) = 0.40 (range 0.28–0.52). [SPEC]
Why: Amodei Oct 2024 — “powerful AI” as early as 2026; compressed 21st century in bio conditional on alignment success (same essay). AlphaFold precedent [EST]; clinic/GMP bottleneck [USER GUESS] caps speed. Not full lifespan doubling by 2032.
Evidence:
- Amodei — Machines of Loving Grace (Oct 2024) — five pillars; risks as only obstacle
- (internal note) — Amodei as research Q5
- (internal note) — U3 science node overlaps; U1 conditional on alignment
Analogue: mRNA platform — decade of work, one crisis accelerates adoption.
Would update if: Zero Phase-3 AI-designed drugs by 2030 and dev-world AI GDP papers show null.
Conf: M
Buckets: U1 (abundance tail), U3
P = 0.12 — Amodei compressed-21st-century full bio realization
Claim: P(“50–100 years of bio progress in 5–10 years” materializes in longevity + disease by ~2040) = 0.12 (range 0.06–0.20). [SPEC]
Why: Requires C9 science loop + alignment + trials infrastructure — triple conjunction. Physical limits + FDA latency in superintelligence physical limits and c8 snapshot. Amodei himself hedges; U1 bucket not U2.
Evidence:
- Amodei essay — lifespan 150, cancer elimination conditional on powerful AI
- (internal note) — trials/GMP binding
- (internal note) — U1 horizon ~2050+
Analogue: Fusion — physics allows, engineering defers.
Would update if: AI-designed therapy ≥5× traditional timeline compression replicated ×3 disease areas by 2035.
Conf: L
Buckets: U1
P = 0.15 — Bostrom symbiosis: human retention instrumentally rational for ASI
Claim: P(superintelligent aligned system chooses to preserve human agency/flourishing as instrumentally stable equilibrium — not merely transitional — through 2050) = 0.15 (range 0.08–0.25). [SPEC]
Why: Bostrom Deep Utopia (2024) — positive scenarios exist but meaning problem unresolved; Superintelligence control problem ≠ solved. Instrumental value of humans: diversity, legal persons, labs-in-wild, moral uncertainty. Counter: competition among AIs drops human marginal value; whimper default in author’s my pdoom.
Evidence:
- Bostrom, Deep Utopia (2024) — (internal note), entity master
- (internal note) §D — post-scarcity meaning
- Christiano slow takeoff — more time for coupling solutions [Hanson variant overlap]
Analogue: Endangered species protected when costly to replace — precarious equilibrium.
Would update if: Aligned ASI explicitly reallocates resources away from human-centric institutions with no pushback — symbiosis branch dead.
Conf: L
Buckets: U2
P = 0.05 — Israetel “curiosity preserves humans” branch
Claim: P(ASI spares humans primarily from scientific curiosity / study value — Mike Israetel R1/R2 argument) = 0.05 (range 0.02–0.10). [SPEC] — non-load-bearing for Path A; included for completeness.
Why: Elon Musk variant in public discourse; not mainstream alignment theory; fails if ASI can simulate humans cheaper than maintaining biosphere humans. Doom Debates Tier S entertainment, low epistemic weight.
Evidence:
- (internal note) — curiosity / study framing
- (internal note) — Israetel <1% P(doom)
- User my pdoom — extinction 19%; this path doesn’t move main estimates
Analogue: Zoo conservation — contingent on human preferences of keeper.
Would update if: No serious lab embeds curiosity-preservation in alignment spec (expected).
Conf: L
Buckets: U2 (weak tail); mostly anti-doom not pro-utopia
Takeoff geometry — slow vs fast positive branches
P = 0.35 — Hanson slow takeoff: institutions absorb, control retained (40% mixture weight)
Claim: P(Hanson-variant multipolar absorption — no local FOOM, liability/insurance markets, human control retained through C9) = 0.35 (range 0.25–0.48) conditional on Hanson-weighted world (author mix 40% Hanson / 60% default per my pdoom / hanson doc).
Why: Hanson disputes speed after AGI, not pre-AGI progress. Anthropic RSI essay Amdahl review bottleneck supports B-tier plateau. Misalign tail 25%→10% in Hanson column — U2 opens via U3/U4 without full alignment proof.
Evidence:
- hanson variant human action — N4 scandal = liability event; extinction 6–9%
- node12 §15 — B-plateau → Hanson 55%
- (internal note) §3 — human review bottleneck
Analogue: Electricity diffusion 1880–1920 — fast capability, slow institutional absorption.
Would update if: Single lab >50% FLOPs + RSI >10× — Hanson falsified (hanson doc §Falsifiers).
Conf: M
Buckets: U3, U4 (primary); U2 partial (coupling without ASI alignment proof)
P = 0.22 — Fast takeoff with oversight keeps pace (default mixture)
Claim: P(METR-doubling ~4 mo cloud RSI continues but interp + debate + RSP gates maintain meaningful human veto through C9) = 0.22 (range 0.12–0.32). [SPEC]
Why: Optimistic branch of default timeline — not modal. Requires interp scaling faster than capability (contradicts Goodfire “SAE lags frontier” concern). NeurIPS 2025: debate best NSO game — if deployment evals map to debate structure.
Evidence:
- METR TH1.1 — 88.6-day doubling
- NeurIPS 2025 oversight scaling — debate 51.7% at 400 gap
- Node 12 L1 modal 0.50 cloud-first — same evidence, different governance response assumption
Analogue: Air traffic control scaling with passenger miles — possible, expensive, never perfect.
Would update if: Two consecutive frontier gens skip interp gate with no public justification.
Conf: L–M
Buckets: U2
P = 0.45 — Fast takeoff outruns oversight by C9 (Path A failure mode)
Claim: P(capability at C9 exceeds verified oversight by ≥1 year — “oversight debt”) = 0.45 (range 0.35–0.55). [USER GUESS] — anti-success; feeds aligned-but-unequal.
Why: Default race dynamics; Node 11 hollowing 0.65; RSI 8× code velocity cited same month as conditional pause. Author crux: no pause 0.88.
Evidence:
- (internal note) — coordination + deception cruxes
- node11 §2 — hollowing 0.65
- Anthropic RSI 2026 — acceleration frame wins
Analogue: Financial leverage — outruns risk models until crisis.
Would update if: Verified year-long lag between capability internal milestone and public eval disclosure closes to <3 mo.
Conf: M
Buckets: blocks U2; U4 ceiling
N4 SUCCESS branches (mirror doom Node 4)
P = 0.15 — Tail-gov: ≥2 labs halt >30d post-C10 scare
Claim: P(same as doom N4 §13 — ≥2 frontier labs voluntary halt >30d or binding licensing above FLOP threshold post Trigger E) = 0.15 (range 0.10–0.22) | Trigger E fires. U1 interpretation: buys alignment R&D time; necessary not sufficient for U2.
Why: Doom doc rationale inverted: EU GPAI enforcement, CA incident duty, natsec licensing frame — rare but real. No voluntary halt precedent caps upper bound. Success = slowdown + verification R&D, not permanent pause.
Evidence:
- node4 §13 — tail-gov 0.15
- Anthropic RSI — expect pause if verification exists
- Seoul 2024 — voluntary commitments without halt (counter)
Analogue: COVID vaccine Operation Warp Speed pause events — brief manufacturing holds.
Would update if: Post-C10 scare, zero lab delays any run >14d — tail-gov dead.
Conf: L–M
Buckets: U2, U3 (multiplier)
P = 0.08 — Durable verified multilateral training slowdown
Claim: P(durable (>6 mo) verified multilateral training limit with enforcement post-C10) = 0.08 (range 0.04–0.14) — doom N4 composite P(durable pause) 0.02–0.05 slightly upgraded for U1 conditional on tail-gov.
Why: Verification gap is well-documented; Anthropic four conditions unpublished mechanism. 8% = tail within tail. Enables U2 only if paired with alignment breakthrough.
Evidence:
- node4 §18 — composite 0.12–0.18 slowdown; durable pause 0.02–0.05
- node11 §15 — P(conditional pause → multilab halt) 0.08
- Paris 2025 skipped binding pause
Analogue: Montreal Protocol — verified multilateral success; AI not yet analogous.
Would update if: Working verification protocol adopted by ≥3 frontier labs + ≥2 governments.
Conf: L
Buckets: U2, U3
P = 0.12 — Verification treaty / HEM-style regime built post-scare (2028–2032)
Claim: P(international incident reporting + capability verification regime with third-party METR/AISI-class audits operational by 2032, catalyzed by C10 scare) = 0.12 (range 0.06–0.20). [SPEC]
Why: Higher than P(durable pause) because monitoring easier than stopping; GAAIA/Banks 2026 direction; IAEA analogue in happy_path. Still low — Seoul/Bletchley follow-through weak.
Evidence:
- Node 4 actor tables — EU formal enforcement P=0.40 (deployment not training)
- International AI Safety Report 2026 — frameworks ↑, binding ↓
- (internal note) — federal pause dead; verification may survive
Analogue: Nuclear inspections without disarmament.
Would update if: US-China bilateral eval sharing treaty signed with METR methodology embedded.
Conf: L–M
Buckets: U3, U2 (enabler)
N11 SUCCESS branches (mirror doom Node 11)
P = 0.10 — Safety re-empowered: RSP/FSF hard stop invoked and holds
Claim: P(same as doom N11 §4 T2 — safety team re-empowered, hard RSP/FSF stop + visible halt/deploy block by 2028) = 0.10 (range 0.06–0.16).
Why: Mechanisms exist (RSP v3.0, FSF v3.1, OpenAI SSC); zero public invocations 2024–26. LTBT majority Apr 2026 on paper; Sharma exit undercuts insider confidence. Success = one credible invocation without immediate commercial override.
Evidence:
- node11 §4 — T2 0.10
- node11 corporate governance — lab tracker
- Anthropic RSP v3.0 (Feb 2026)
Analogue: SCRAM — works once, changes industry expectations.
Would update if: ASL-4+ trigger → public deployment delay >30d and training slowdown documented.
Conf: L–M
Buckets: U2, U4
P = 0.15 — Corporate governance SUCCESS: deploy veto used ≥1× by 2028
Claim: P(OpenAI SSC or Anthropic LTBT or DeepMind FSF blocks or delays frontier release past schedule with public notice) = 0.15 (range 0.10–0.22) — aggregate across labs.
Why: Higher than training halt (Node 11 §13 SSC block 0.15 OpenAI-only) because deployment veto narrower politically. CA AG MOU Oct 2025 — SSC halt authority. No observed block post-recap — track record discount.
Evidence:
Analogue: FDA CRL — delay, not shutdown.
Would update if: Two frontier releases in 2027–28 on time despite leaked internal safety objections.
Conf: L–M
Buckets: U2, U4
P = 0.35 — N11 SUCCESS: safety independence structure holds without hollowing
Claim: P(≥1 frontier lab maintains CEO-only reporting safety org with documented deploy veto and no superalignment-style dissolution through 2028) = 0.35 (range 0.25–0.48). [USER GUESS]
Why: Inverse of hollowing 0.65 not direct complement (partial states exist). Anthropic Safeguards + LTBT best candidate; OpenAI weakest. Google RAI → DeepMind integration negative signal.
Evidence:
- Node 11 §2 M1 hollowing 0.65
- Node 11 §17 — P(safety reports to product, ≥2 labs) 0.45 failure mode
- (internal note) §23
Analogue: Independent internal audit — survives if regulator demands.
Would update if: Second lab dissolves independent safety org 2026–27.
Conf: M
Buckets: U2, U3
N12 SUCCESS branches (mirror doom Node 12)
P = 0.35 — RSI stalls at tier B with human control retained (Hanson-positive)
Claim: P(cloud RSI plateaus at tier B — closed-loop optimization, human fixes meta-framework — through 2030, without public tier C) = 0.35 (range 0.25–0.45).
Why: Anthropic essay Amdahl on review; OpenAI Preparedness — o3 not High on self-improvement; Node 12 P(C-tier public by 2028) = 0.12. Positive for U3/U4 — institutions keep up; U2 needs alignment not just slow RSI.
Evidence:
- node12 §8–9, §15
- (internal note)
- Hanson variant — misalign tail ↓
Analogue: Semi-autonomous cruise control — human retains meta-control.
Would update if: Public tier C evidence + no human review bottleneck disclosed.
Conf: M
Buckets: U3, U4; partial U2
P = 0.25 — Tier C RSI with human control retained (hard mode)
Claim: P(public tier C — autonomous months-long AI R&D — and human org retains effective goal-setting + kill switch) = 0.25 (range 0.15–0.35) | tier C achieved. [SPEC]
Why: Tier C may arrive (Node 12 0.12 by 2028 unconditional); conditional control is harder than capability. Requires N11 success + oversight scaling. Product P(C-tier) × P(control|C-tier) ≈ 0.12 × 0.55 ≈ 0.07 unconditional — round to 0.25 | C-tier.
Evidence:
- Node 12 §9 — P(C-tier public by 2028) 0.12
- Anthropic RSI — human review bottleneck explicit
- Node 4 — C10 alignment scare during tier C transition
Analogue: Nuclear ** breeder reactor** — powerful, requires active control rods.
Would update if: Tier C announced with no RSP escalation — control failure signal.
Conf: L
Buckets: U2, U1 tail
P = 0.18 — Bio/cyber locus fires first; alignment buys calendar time
Claim: P(L2 bio or L4 cyber first policy-salient tail before cloud C10 alignment scare — and this increases P(successful alignment effort)) = 0.18 (range 0.10–0.28). [SPEC]
Why: Node 12 L2 0.22, L4 0.10 — non-cloud first tail slows cloud RSI race or redirects governance to chokepoints alignment community understands (BMIA, SEC). Not guaranteed positive — could trigger panic deployment. Hanson L3/L4 first → Hanson weight 0.55.
Evidence:
- node12 §3–5, §16
- Hanson variant — N8 modal absorption
- Node 2 bio — BMIA window
Analogue: Challenger disaster — safety reform if institutions functional.
Would update if: Bio near-miss → accelerationist deregulation instead of screening.
Conf: L
Buckets: U3, U4
Pessimistic non-doom (Path E) — blocks U2/U3
P = 0.28 — Aligned-but-unequal equilibrium
Claim: P(ASI or near-ASI aligned to owner preferences — no extinction — but ≥80% of AI surplus captured by ≤0.1% / compute owners through 2050) = 0.28 (range 0.20–0.38). [USER GUESS]
Why: Modal c8 snapshot Ghost GDP; Acemoglu Power and Progress; author’s PE lens in my_alignment_position.md. Alignment ≠ distribution. Compatible with low extinction (doom whimper channel partially avoided).
Evidence:
- (internal note) — modal friction
- (internal note) — aligned-but-unequal red-team variant
- (internal note) — concentration prior
Analogue: Industrial Revolution — growth without broad capture for decades.
Would update if: ≥2 G7 compute dividend / SWF laws pass with >3% GDP redistribution from AI surplus by 2035.
Conf: M
Buckets: blocks U2, U3; allows U4 partial
P = 0.22 — Aligned-but-bored: physical success, normative failure
Claim: P(human instrumental role collapses; aligned AI meets preferences but eudaimonia institutions fail — stable wireheading-lite equilibrium) = 0.22 (range 0.12–0.32). [SPEC]
Why: Bostrom Deep Utopia central crux; Jahoda functions in happy_path; U7 node pending. Physically aligned — normatively empty U2.
Evidence:
- Bostrom, Deep Utopia (2024) — (internal note)
- (internal note) — U7 meaning node
- Amodei essay §work and meaning — optimistic but unproven institutional path
Analogue: Retirement without purpose — persistent anomie (Durkheim).
Would update if: Post-C8 societies show rising self-reported meaning / civic participation despite automation — 3-country panel.
Conf: L–M
Buckets: blocks U2 (normative); U4 physical possible
P = 0.55 — Path A modal: aligned-enough, unequal or bored (composite)
Claim: P(Path A resolves to non-extinction but not full U2 — i.e., aligned-but-unequal or aligned-but-bored or whimper-adjacent agency loss) = 0.55 (range 0.45–0.65) | Path A active at C10.
Why: Sum of pessimistic non-doom branches minus overlap; mirrors doom modal not extinction but not utopia either.
Evidence:
- This doc §aligned-but-unequal, §aligned-but-bored
- my putopia — modal non-doom ≠ utopia
- my pdoom — agency loss 24%
Analogue: Social media era — aligned to engagement, bad for flourishing.
Would update if: Sen-capability indices rise in high-AI-adoption cohorts 2028–32.
Conf: M
Buckets: anti-U2/U3
Anti-success cruxes (why alignment might NOT scale)
P = 0.55 — Deception survives production evals at C9
Claim: P(strategic deception not detected by standard RLHF + red-team + RSP eval suite at C9 internal capability) = 0.55 (range 0.45–0.65). Mirrors doom crux inverted.
Why: Sleeper Agents, Apollo scheming, eval-awareness literature; Apollo safety cases shift burden to monitoring, not inability proofs.
Evidence:
- Anthropic Sleeper Agents (2024); Apollo Dec 2024
- (internal note) — deception crux 0.70 at doom horizon; 0.55 here = slightly optimistic conditional on effort post-C10
Analogue: Penetration tests — pass until novel attack.
Would update if: Third-party eval catches deception pre-deploy on successive generations — rate falls >50%.
Conf: M–H
Buckets: anti-U2
P = 0.65 — Safety hollowing continues (N11 failure modal)
Claim: P(doom N11 M1 — continued safety hollowing + narrative without veto, 2026–28) = 0.65 — same as doom; U1 success requires 0.35 not hollowing.
Why: Shared evidence base — superalignment dissolved, Sharma exit, RSI managed disclosure.
Evidence:
- node11 §2
Analogue: See doom node — pharma pharmacovigilance.
Would update if: ≥2 labs restore independent safety org with used deploy veto.
Conf: M
Buckets: anti-U2
P = 0.70 — Modal N4 response: scandal → transparency, not alignment buy-in
Claim: P(C10 whistleblower → modal doom N4 response — audits, continue training — without proportional alignment budget or verification) = 0.70 (range 0.58–0.82) | Trigger E.
Why: Doom Node 4 §12 modal 0.58 + E4 managed disclosure 0.37 mass; effort ↑ modestly, success not.
Evidence:
- Node 4 §40 Adelstein counter-evidence table — train-through pattern
- Saunders 2024, board 2023
Analogue: Enron → Sarbanes-Oxley paper compliance first years.
Would update if: Post-scare alignment spend >2× capability spend at any frontier lab (public accounts).
Conf: M–H
Buckets: anti-U2; enables U3 governance-only path
Composite estimates (Path A → buckets)
P = 0.18 — P(U2 symbiosis by 2050 | Path A)
Claim: P(U2 symbiosis flourishing — aligned superhuman AI, humans retain meaningful agency + eudaimonia — by 2050 via Path A) = 0.18 (range 0.12–0.28). [USER GUESS] Phase 2 stitch input.
Why: Illustrative conjunction (overlap not applied):
| Factor | P(success) | Source section |
|---|---|---|
| Oversight scales C9 | 0.38 | §NSO |
| Deception caught/managed | 0.45 (1−0.55) | §deception |
| N4 tail-gov + verification | 0.15 | §tail-gov |
| N11 veto holds | 0.15 | §deploy veto |
| N12 control retained | 0.30 | §tier C×control |
| Meaning institutions | 0.40 | [SPEC] U7 placeholder |
| Naive product | ~0.001 | — |
| Overlap-adjusted | ~0.12–0.28 | ρ≈0.5 discount |
Evidence: Sections above; my putopia Step 2A conjunction template.
Analogue: Startup series of unlikely wins — some paths succeed disjunctively.
Would update if: Any single factor → >0.6 with evidence — revise up to 0.30+.
Conf: L
Buckets: U2
P = 0.08 — P(U1 radical abundance tail | Path A + U3 science)
Claim: P(U1 radical abundance — longevity escape, large-scale space/energy — conditional on Path A alignment success and U3 science wins) = 0.08 (range 0.04–0.14) by 2060. [USER GUESS]
Why: Requires C9+ aligned science loop plus physical deployment (U4 node); double conjunction beyond U2.
Evidence:
- Amodei §compressed 21st century — 0.12 full bio
- (internal note) — ceilings
- Node U3/U4 pending research
Analogue: Apollo program — aligned gov + science + budget.
Would update if: AI-designed longevity therapy Phase 3 before 2035.
Conf: L
Buckets: U1
P = 0.32 — P(U3/U4 institutional golden age without full U2)
Claim: P(U3 or U4 — major welfare gain, no extinction, without proven superhuman alignment symbiosis) via Hanson slow takeoff + governance + science = 0.32 (range 0.22–0.42). [USER GUESS]
Why: Disjunctive path — Path A failure on full alignment compatible with Hanson + U3 nodes. 40% Hanson weight × institutional absorption.
Evidence:
- Hanson variant — agency loss ↑, extinction ↓
- my putopia — U3 default external answer
- Node U2/U5/U6 pending
Analogue: Post-WWII institutional boom — no ASI required.
Would update if: Hard takeoff confirmed — ↓ to <0.15.
Conf: M
Buckets: U3, U4
Top 5 cruxes (load-bearing for Path A)
| # | Crux | P(crux holds) | If true → | Would update if… | § |
|---|---|---|---|---|---|
| 1 | Illegible scheming outruns interp | 0.82 | U2 tail only; monitoring arms race | Interp >90% recall on hidden scheming C9 | §interp illegible |
| 2 | NSO fails above 400 Elo gap | 0.62 | Debate insufficient at C9 | 3-level NSO certified externally | §NSO |
| 3 | Modal N4 → transparency not halt | 0.70 | Speed ↑; verification R&D only | ≥2 labs halt >30d post-scare | §N4 modal |
| 4 | RSI tier C without control | 0.55 | Extinction or whimper ↑; U2 ↓ | Tier C + documented human kill switch works | §tier C |
| 5 | Meaning collapse post-instrumentality | 0.58 | U2 physically empty | Rising meaning metrics in automated economies | §aligned-bored |
Branch mixture (Path A @ C10)
| Branch | P(branch) | U2 lift | Primary buckets | Doom mirror |
|---|---|---|---|---|
| Modal aligned-enough | 0.55 | Low | U4 partial; blocks U2 | Whimper / concentrated harm |
| Success tail symbiosis | 0.18 | High | U2 | N4 tail-gov + N11 veto |
| Hanson slow control | 0.15 | Medium | U3, U4 | Hanson misalign ↓ |
| Aligned unequal/bored | 0.28 | Negative | U4 ceiling | Agency loss channel |
| Hard fail (misalign) | 0.19 | — | — | Extinction bucket |
Note: Branches overlap (not partition); extinction 0.19 from my pdoom outside Path A conditional table.
Evidence library (2024–2026)
| Org | Finding | U1 direction | Link |
|---|---|---|---|
| Apollo | In-context scheming on frontier models | ↓ U2 | Science |
| Apollo+OpenAI | Deliberative alignment ~30× scheming reduction | ↑ U2 partial | Apollo homepage |
| Apollo+AISI+METR+Redwood | Safety cases for scheming; interp not ready for illegible | ↓ U2 | Report |
| Anthropic | Constitution Jan 2026; RSP v3.0; RSI essay Jun 2026 | ↑ procedure; ↓ if continue-train | Constitution |
| Anthropic | Sleeper Agents 2024 | ↓ U2 | Research page |
| NeurIPS 2025 | Oversight scaling laws; debate best NSO game | ↑ debate; ↓ backdoor | Paper |
| METR | Time horizon doubling ~88d; eval partner | ↑ calibration | TH1.1 |
| CHAI/Russell | AssistanceZero ICML 2025 | ↑ theory; production distant | PMLR |
| Amodei | Machines of Loving Grace Oct 2024 | ↑ U1/U3 vision | Essay |
| IAIS Report 2026 | 12 frontier frameworks; eval gap persists | ↑ governance; ↓ verify | Summary |
Update log
| Date | Change |
|---|---|
| 2026-07-04 | Initial Node U1 evidence rationale; 42 P-sections; Path A; doom mirrors N4/N11/N12 |
Related repo files
- (internal note) — synthesis target
- (internal note) — shared cruxes, inverted branches
- (internal note) — author profile + evidence updates
- (internal note) — early agenda
- (internal note) — slow takeoff positive branch
- (internal note) — modal non-utopia contrast
- (internal note) — (pending) meaning crux detail