← Back to all writing

AI futures evidence — U6: multilateral beneficial AI

July 5, 2026

Futures index · 中文 · Main forecast

Each section: Claim · Why · Evidence · Analogue · Would update if · Conf (H/M/L).


Parent: Utopia timeline
Mirror (doom failure branches): node9 (N9 multilateral), node3 (N3 T2 treaty / US–China)
Modal baseline (NOT utopia): c8 society snapshot by ci
Date: 2026-07-04
Format: Claim | Why | Evidence | Analogue | Would update if | Conf

Each section documents one probability for U6 — multilateral beneficial-AI coordination: summit follow-through, deployment standards, verification R&D, and credible (not empty-shell) international AI governance. Probabilities are subjective elicitation unless noted as derived from N9/N3.

Utopia buckets: U3 multiplier — international standards reduce friction for science dividend and institutional golden age; does not directly unlock U1/U2 without U1/U5.


1. Node definition — beneficial coordination ≠ pause

Claim: U6 = P(by @C10, multilateral AI governance produces materially beneficial outcomes — shared evals, deployment incident norms, verification R&D funding, Global South compute access — with ≥1 binding enforcement path or verified pilot, without requiring global training pause).

Why: N9 modal (0.58) = forum talk without enforcement — anti-U6. U6 success tails = N9 branches T1 partial deployment (0.18), T3 verified limits (0.05), plus positive interpretation of India/UN inclusion. Anthropic conditional pause with verification is U6 success endpoint, not current expectation.

Evidence:

  • Utopia timeline §U6
  • N9 parent — branch M1/T1/T2/T3
  • N3 §17 T2 — real treaty tail 0.04–0.08

Analogue: IAEA safeguards — verification without ending nuclear power globally.

Would update if: Verified training-limit treaty ratified → U6 merges with pause track; revise node boundary.

Conf: H


2. Branch overview — U6 mixture @C10

Claim:

BranchPDescriptionN9 mirror
U6-M (forum + soft law)0.50Bletchley→Seoul→Paris→India→Geneva continuity; AISI network; no binding verificationN9 M1 0.58
U6-T (partial enforcement)0.32EU GPAI + CoE treaty + Korea/Japan deployment fines + US–China guardrails + HEM pilotN9 T1 0.18 + uplift
U6-S (strong coordination)0.10Verified training limits or IAEA-analog with inspection rights + Anthropic pause activatedN9 T3 + N3 T2
U6-X (empty-shell trap)0.0890+ signatories, zero enforcement — false confidenceN9 T2 0.10

Why: U6-T > N9 T1 alone because deployment enforcement already partially live (EU AI Act Aug 2025/2026). U6-S stays thin tail. Sum ≈1.0.

Evidence:

  • N9 §2 branch overview
  • N9 §28 treaty composite table

Analogue: Climate COP — ambition declarations (M) vs Paris Agreement Article 13 transparency (T) vs binding cap-and-trade (S, rare).

Would update if: Geneva 2027 mandates weights-audit protocol → U6-S .

Conf: M


3. P = 0.22 — Composite U6 “beneficial enough” @C10

Claim: P(by @C10, ≥3 of 5: (1) AISI-network eval sharing operational, (2) EU GPAI enforcement with ≥2 actions, (3) US–China guardrails track with joint standard, (4) UN Scientific Panel annual cycle credible, (5) HEM or compute-attestation pilot with ≥2 countries) = 0.22 (range 0.15–0.30).

Why: Discounted joint — correlated with U5 EU/US domestic wins. Beneficial enough for U3 multiplier, not utopia. Modal c8 includes subset (EU enforce + AISI) at ~0.15–0.18 — U6 composite slightly above modal governance alone.

Evidence:

  • Component §9–§22
  • c8 C6–C8 governance rows

Analogue: WTO TBT — mutual recognition without tariff zero.

Would update if: Phase 3 correlation_matrix_positive.md refines overlap with U5.

Conf: M


4. Cross-ref c8 — modal multilateral = talk + EU enforce, still no pause

Claim: c8 modal multilateral posture: summit diplomacy continues; EU faster than US on deployment; US–China guardrails talk not treaty; P(no pause) 0.82–0.84 unchanged.

Why: c8 §Governance row — “EU GPAI enforcement; no training FLOP cap; durable global pause P <2%.” U6-M is c8 modal, not utopia success.

Evidence:

Analogue: UNFCCC process — decades of COPs, emissions continue.

Would update if: c9 utopia snapshot doc quantifies U6 uplift per Ci — calibrate there.

Conf: H


5. Bletchley 2023 — institutional seed (success foundation)

Claim: Bletchley Nov 2023 established durable institutional vocabulary (frontier AI, evals, AISI network) with P(institutional legacy persists through 2028) = 0.85.

Why: Seoul, Paris, India, Geneva all cite Bletchley lineage; 28 countries → 91 signatories trajectory. Success = agenda persistence even when binding content narrows.

Evidence:

  • (internal note)
  • (internal note) §2.8
  • UK gov Bletchley communique

Analogue: Bretton Woods conference — institutions outlast initial ambitions.

Would update if: Major summit explicitly repudiates Bletchley safety frame → .

Conf: M–H


6. Seoul 2024 — AISI network + Frontier AI Safety Commitments

Claim: P(≥8 countries actively share frontier eval methodologies via AISI network by 2028) = 0.72; P(China signs ministerial declaration) = 0.15 (low — success = corporate Zhipu-style commitments).

Why: Seoul produced AISI network (10 countries + EU); ministerial 27+EU without China prefigured Paris split. Beneficial = eval sharing, not pause.

Evidence:

Analogue: IPPC — scientific coordination without emissions cap.

Would update if: China co-signs next Seoul-style ministerial with US → 0.35+.

Conf: M


7. Paris 2025 — fragmentation with partial success

Claim: Paris AI Action Summit (Feb 2025): ~60 signatories including China, India, UAE; US/UK unsigned — P(this pattern continues through 2028) = 0.70; P(beneficial outcome despite US absence: Global South agenda-setting + investment pledges) = 0.55.

Why: US absence hurts verification treaty; helps inclusive/sustainable frame aligning India/UAE with access not pause — U6 partial win for U3 Global South science access.

Evidence:

Analogue: Kyoto without US — regime continues, weaker but not empty.

Would update if: Second Trump term US signs successor with verification annex → branch break.

Conf: M–H


8. India AI Impact 2026 — New Delhi Declaration (91 signatories)

Claim: P(New Delhi Declaration ≥100 follow-on signatories by 2028 with operational compute-access or sovereign-AI pledges) = 0.60; P(anti-freeze frame persists) = 0.85.

Why: 88→91 signatories Feb 2026; US and China signed; “AI for All” / sovereign AI — beneficial for inclusion, hostile to pause. U6 success redefined as access + standards, not halt.

Evidence:

Analogue: NIEO 1970s technology transfer rhetoric — partially real infra (India Stack), not capacity restraint.

Would update if: Declaration amended with verifiable compute ceiling → anti-freeze ; U6-S .

Conf: M–H


9. UN Global Dialogue Geneva Jul 2026 — follow-through

Claim: P(Jul 6–7 2026 Geneva session produces credible roadmap to 2027 NY session with ≥120 country participation) = 0.65; P(mandatory enforcement outcome) = 0.05.

Why: GDC mandate = non-regulatory inclusive forum; 1,000+ written submissions; Scientific Panel preliminary report Jul 1. Success = process legitimacy + science feed, not treaty.

Evidence:

Analogue: WSIS 2003–2005 — decade process, soft outcomes first.

Would update if: Outcome doc includes mandatory national reporting on >10²⁵ FLOP runs → U6-T .

Conf: M


10. UN Independent Scientific Panel — annual evidence cycle

Claim: P(Panel delivers ≥2 credible annual reports (2026 preliminary + 2027 full) used in ≥50 national AI policies) = 0.58.

Why: Bengio + Ressa co-chairs; 40 experts from 2,600 candidates; preliminary report Jul 2026 warning: safeguards lag capabilities. Beneficial = shared epistemic baseline — prerequisite for any verification treaty.

Evidence:

Analogue: IPCC — science coordination precedes (weak) policy.

Would update if: Panel disbanded or politicized before 2027 report → .

Conf: M


11. P = 0.10 — Empty-shell treaty (dangerous failure mode)

Claim: P(global signed AI instrument ≥80 parties with no verification/enforcement body operational by 2030, cited as “governance” by labs) = 0.10 — N9 T2 inverted as U6 failure.

Why: Paris + New Delhi already match profile — broad signatories, voluntary pillars. Danger: false confidence accelerates race (N9 p(doom) +0.5–1pp). U6 must track separately from success.

Evidence:

Analogue: Paris Agreement early NDCs — pledged without enforced compliance.

Would update if: Signatories fund IAEA-analog with inspection rights and US joins → reclassify as U6-S.

Conf: M


12. P = 0.18 — Partial deployment standards with real enforcement

Claim: P(binding deployment rules — incident reporting, documentation, market withdrawal — without training cap, enforced on frontier providers globally) = 0.18 — N9 T1 success tail.

Why: EU AI Act live; Council of Europe CETS 225 (2024); UK AISI evals; Korea AI Basic Act trajectory. 18% = compound path where Brussels effect + CoE + AISI network align.

Evidence:

Analogue: Basel III — heavy reporting, no ban on banking.

Would update if: US federal law preempts extraterritorial EU duties → ; Korea fines US frontier lab → .

Conf: M


13. P = 0.05 — Verified training-limit treaty by 2030

Claim: P(bilateral or multilateral training FLOP ceiling + verification treaty by 2030) = 0.05 (range 0.03–0.08) — N9 §13 / N3 T2.

Why: HEM immature; US–China agenda mismatch; India anti-freeze. 5% = extreme tail — U6-S core. Requires crisis + verification breakthrough.

Evidence:

  • N9 §13; N3 §17 T2
  • arxiv 2604.04712 — treaty-grade metering immature (Node 4 §19)

Analogue: INF Treaty 1987 — parity + verification tech + political window.

Would update if: US–China guardrails publish pilot HEM with public results → to 0.10–0.15.

Conf: M–H


14. IAEA / NPT analogue — feasibility assessment

Claim: P(IAEA-analog AI agency with inspection rights on training clusters funded and chartered by 2030) = 0.08; P(full operational verification) = 0.03.

Why: NPT succeeded on fissile material chokepoint + geopolitical window; AI lacks (1) agreed hazard material, (2) US–CN parity bargain, (3) HEM maturity. Analogue informs U6-S shape, not base rate.

Evidence:

  • N9 §IAEA references; research_international_ai_governance_platforms.md
  • N3 §21 — weights verification immature
  • Hinton / Global Call for AI Red Lines 2026 — aspirational

Analogue: IAEA 1957 — years after Hiroshima + Eisenhower Atoms for Peace.

Would update if: UN GA emergency resolution creates AI verification fund with US+CN co-sponsorship → .

Conf: L


15. P = 0.50 — US–China guardrails talks continue (not pause)

Claim: P(formal US–China AI dialogue restarts and continues through 2028 on guardrails/misuse) = 0.50; P(mutual training freeze | talks) <0.02.

Why: May 2026 Beijing summit restart; Bessent non-state-actor frame; Track II >8 meetings vs Track I 1 session ever. Talks decouple from pause — CN conditions on chip controls.

Evidence:

  • (internal note)
  • N9 §14; N3 §20

Analogue: US–SU hotline — crisis comms, not disarmament.

Would update if: Second Geneva session publishes joint safety standard → U6-T .

Conf: M


16. P = 0.06 — US–China verification R&D joint pilot (N3 T2 success slice)

Claim: P(US + China publicly fund or announce joint HEM / compute-attestation R&D pilot by 2028) = 0.06 — subset of N3 T2 tail.

Why: T2 = training limits + partial verification; pilot alone = U6-T win without cap. Theft crisis usually accelerates race (T3 0.15–0.25) more than enables trust — low base rate.

Evidence:

  • N3 §17 T2 P=0.04–0.08
  • N3 §21 — verification gap
  • AI 2027 Appendix S — HEM / compute pause concepts

Analogue: US–RU lab-to-lab cooperation 1990s — narrow trust-building.

Would update if: Post-Geneva 2026 bilateral MOU on compute verification → to ≥0.12.

Conf: L


17. HEM / compute metering immaturity — binding constraint

Claim: P(treaty-grade HEM operationally ready for frontier training verification by 2030) = 0.15 (range 0.10–0.22).

Why: FlexHEG academic; export-control attestation partial; no deployed mandatory remote attestation on H100 clusters. Playbook: 3–5 yr medium horizon. Blocks U6-S until resolved.

Evidence:

  • N3 §21; Node 7 §HEM P=0.12 pilot
  • (internal note) — prerequisite #4

Analogue: Nuclear test detection — years of seismic network build.

Would update if: Major cloud mandates HEM for gov contracts → .

Conf: M


18. Anthropic conditional pause with verification — positive path

Claim: P(Anthropic actually halts frontier training >30 days citing verified multilateral agreement | such agreement exists) = 0.70; P(agreement exists by 2030) = 0.06; Unconditional P(halt) ≈ 0.04.

Why: Jun 2026 RSI policy essay — pause only if multilateral verification exists; doesn’t exist. Positive path = U6-S enables lab credible halt — not current modal. Rhetoric expect, not commit.

Evidence:

  • (internal note)
  • N3 §15 — Anthropic halt 0.05–0.08 unilateral
  • N9 §2 — Anthropic RSI Jun 2026

Analogue: Pharmaceutical trial voluntary hold — real when FDA + verified data exist.

Would update if: ≥2 labs halt >90 days citing New Delhi or Geneva instrument → conditional validated.

Conf: M


19. Global South inclusion — 118-country gap

Claim: P(≥118 UN member states meaningfully participate in AI governance forum — speaking slots, written submissions, compute-access pledges — by 2028) = 0.75; P(binding voice on verification design) = 0.20.

Why: Geneva “every country a seat”; India 91 signatories; Global South underrepresented in Bletchley ministerial core. Inclusion successpause — often anti-freeze.

Evidence:

Analogue: WTO Doha development round — inclusion without power shift.

Would update if: UN Dialogue weighted voting on AI rules → binding voice .

Conf: M


20. Global South anti-freeze coalition — U6 tension

Claim: P(India-led bloc blocks training moratorium language in any multilateral instrument through 2030) = 0.80.

Why: New Delhi seven chakras — access, growth, energy; PM Modi “Sarvajan Hitaya”; sovereign AI as development right. Beneficial U6 must work with this constraint.

Evidence:

  • N9 §6, §21 — P(anti-freeze)=0.80
  • Node 4 §27 — DeepSeek → accelerate not pause

Analogue: G77 on climate — common but differentiated responsibilities.

Would update if: India hosts pause summit coalition → falsify.

Conf: M–H


21. P = 0.75 — UK–Australia AISI eval sharing (Five Eyes lane)

Claim: P(UK–Australia May 2026 MoU produces ≥2 shared classified or dual-use frontier eval reports used in policy by 2028) = 0.75.

Why: Concrete bilateral beneficial coordination below treaty threshold; cyber/natsec emphasis post-AISI rebrand. U6-T building block.

Evidence:

Analogue: Five Eyes SIGINT sharing — deep cooperation without public treaty.

Would update if: MoU never produces shared public or leaked eval product → .

Conf: M


22. Council of Europe CETS 225 — treaty layer

Claim: P(≥35 parties ratify AI Framework Convention with enforceable human-rights oversight on high-risk AI by 2028) = 0.45.

Why: Opened Sep 2024; first binding multilateral AI treaty; overlaps EU but extends to UK, Canada, etc. Deployment-centric — U6-T compatible.

Evidence:

Analogue: European Convention on Human Rights — regional binding layer.

Would update if: US signs CETS 225 → major.

Conf: M


23. Korea / Japan / Brazil — deployment enforcement wave

Claim: P(≥1 of Korea AI Basic Act, Japan AI Promotion Act, Brazil Marco Legal AI produces fine ≥$1M on foreign frontier provider by 2029) = 0.28.

Why: N9 falsifiers watchlist — Korea fines US lab T3; Japan MATCH alignment; Brazil Chamber vote uncertain. Compound Brussels effect.

Evidence:

  • N9 §32 falsifiers — Korea AI Basic Act
  • N9 §38–§42 regional rows

Analogue: GDPR fines on US tech — extraterritorial enforcement.

Would update if: All three fail enactment → to <0.12.

Conf: L–M


24. UAE / Gulf — third-pole beneficial coordination

Claim: P(Gulf states coordinate sovereign compute + Paris/New Delhi pledges into regional eval hub used by Global South partners) = 0.40.

Why: UAE Paris signatory; G42 Stargate; MGX/Mistral — beneficial = compute access + standards export, not pause. Increases global FLOPs (race channel) while improving local governance capacity.

Evidence:

  • N9 §24 — UAE P(Paris signatory)=0.95
  • Gulf AI policy architecture 2026 sources in N9

Analogue: Singapore financial hub — standards + access.

Would update if: EU blocks G42 exports over GPAI → regional hub .

Conf: L–M


25. Correlation with P(no pause) — U6 beneficial ≠ halt

Claim:

U6 branchP(no pause) conditionalDelta vs 0.82 baseline
U6-M forum talk0.84+0.02 (empty rhetoric)
U6-T partial enforcement0.80−0.02 (deployment friction only)
U6-S verified treaty0.55–0.65−0.17–0.27
U6-X empty-shell0.86+0.04 (false confidence)

Why: N9 §29 weighted math; empty-shell worst for alignment timing. Do not multiply marginals — scenario branches.

Evidence:

Analogue: Trade agreements — more rules, commerce continues.

Would update if: U6-S realized → revise my pdoom crux #5.

Conf: M


26. U3 multiplier — international standards reduce science friction

Claim: P(U3 science dividend path | U6-T vs U6-M) = +0.10–0.15 absolute — shared evals + incident norms reduce cross-border research friction and surprise market withdrawals.

Why: U6 enables reproducible safety baselines for biotech/materials AI tools deployed multi-jurisdiction. Conditional on U5 domestic screening — correlated ρ≈0.4.

Evidence:

  • Master doc §U6 — U3 multiplier
  • c8 science bottleneck — trials/GMP dominate; U6 marginal on discovery

Analogue: ICH harmonization — faster drug approval across jurisdictions.

Would update if: node u3 science abundance quantifies joint model.

Conf: L–M


27. Empty-shell as false security — p(doom) channel

Claim: U6-X empty-shell +0.5–1.0pp to P(extinction by 2050) vs counterfactual — policymakers/labs cite signatures while race accelerates.

Why: N9 §30 p(doom) channels; Hinton warning on advisory bodies. Worst U6 branch for timing — worse than honest forum talk (U6-M).

Evidence:

  • N9 §11, §30
  • LessWrong Global Call for AI Red Lines — aspirational 2026 deadline

Analogue: Maginot Line — confidence without coverage.

Would update if: Independent audit shows summit signatories slowed frontier releases → falsify false-security thesis.

Conf: L–M


28. What U6-S would look like if it happened

Claim: U6-S observable bundle by 2030: (1) US–CN ratified side agreement with HEM pilot + declared FLOP notification, (2) UN-chartered verification body with ≥3 inspection events, (3) Anthropic + ≥1 peer 30-day halt after verified trigger, (4) India does not veto — requires ** semiconductor bargain**.

Why: Scenario sketch for falsifiability — not forecast. N3 Slowdown Ending “The Deal” reference class.

Evidence:

  • N3 §17 T2 description
  • AI 2027 Slowdown Ending narrative (scenario device)

Analogue: SALT I — limited, verified, politically contingent.

Would update if: Any 2 of 4 observables met → U6-S to ≥0.25.

Conf: L


29. P(treaty) composite — paper vs enforcement (U6 accounting)

Claim:

MetricPU6 classification
Any signed instrument ≥60 parties0.85Already true — not success alone
Empty-shell (no verification)0.10U6-X failure
Partial deployment enforcement0.18U6-T core
Verified training limits0.05U6-S core
Net beneficial (U6-T or better)0.22–0.32§3 composite

Why: Separates paper from enforcement — N9 §28 master update.

Evidence:

  • N9 §28

Analogue: Double-entry bookkeeping — gross signatories vs net enforcement.

Would update if: Phase 3 synthesis merges with U5 stack probability.

Conf: M


30. Actor table — multilateral success menu

ActorBeneficial outcomePConf
USGuardrails bilaterally; signs deployment standard0.55 / 0.15M
ChinaOpen-weights diplomacy + guardrails talk0.70 / 0.50M
EUGPAI enforcement + CoE alignment0.40 / 0.45M
IndiaSovereign AI + declaration host0.55 / 0.80 anti-freezeM
UK/AUAISI eval sharing0.75M
UNScientific Panel + Dialogue roadmap0.58 / 0.65M
Frontier labsConditional pause if verified0.04 unconditionalM

Evidence: N9 §19–§24 actor tables.


31. Falsifiers (watchlist)

ObservationBranch
US+CN verified training treatyU6-S live
EU no GPAI fines by 2029U6-T
≥2 labs halt >90d citing multilateral dealP(no pause) major
India pause coalition§20 wrong
Geneva 2027 weights-audit mandateU6-S
100+ signatories, zero enforcement eventsU6-X

Conf: M



Update log

DateChange
2026-07-04Initial research pass — 32 sections; Bletchley→Geneva chain + N9/N3 success tails