← Evidence index · 中文 · Main post
Each section: Claim · Why · Evidence · Analogue · Would update if · Conf (H/M/L).
Parent timeline: Shared Ci spine §Node 4 (lines 146–284)
Date: 2026-07-03
Settings: Hybrid time (C); modal + tail branches; tracker pace ~0.70×
Purpose: Load-bearing evidence and reasoning for every probability in Node 4 — Trigger E variants, scenario branches, actor table, composite priors, and P(extinction | E, modal) chain.
1. Executive summary
Node 4 is the alignment-salience human-action node: at capability C9–C10 (superhuman-AI-researcher tier), internal evals flag deception/scheming; a Trigger E variant determines whether this becomes a mainstream whistleblower cycle or stays classified/corporate-managed.
Central claim (modal, P=0.58 given Trigger E fires): 6–10 week media cycle → GAAIA/Banks-style oversight + audits + incident reporting → frontier labs continue training (maybe 10–15% marginal slowdown) → no mandatory pause → conditional multilateral pause rhetoric without verification regime.
Composite prior: P(meaningful training slowdown | whistleblower-class event) ≈ 0.12–0.18; P(durable multilateral pause) ≈ 0.02–0.05.
p(doom) slice: P(extinction by ~2050 from misaligned superintelligence | C10 scare + modal response) ≈ 12–22% (mid ~17%) — Node 4 is the coordination crux for the misalignment bucket; modal path does not materially cut extinction risk.
2. Where this fits in the repo
| File | Role |
|---|---|
| Shared Ci spine | Master timeline; Node 4 summary tables |
| This file | Full evidence rationale for Node 4 probabilities |
| timeline prediction node2 cbrn full | Parallel CBRN tree (uncorrelated with whistleblower path) |
| research_ai_pause_advocacy_playbook.md | Policy mechanics; why pause is dead federally |
| Anthropic_When_AI_Builds_Itself_2026.md | E4 analogue; conditional pause parsing |
| AI_safety_大事记_2022-2026.md | Historical analogue timeline |
| doom_debates_tier_SA_中文讲解.md | Adelstein S1 conjunction; Rob Miles A4 warning-shot crux |
| my pdoom | Calibration target for conditional extinction chain |
3. Node definition — capability anchor C10
Capability C: Frontier internal evals at C9–C10 — superhuman AI-researcher tier (AI 2027 Agent-4 analog; ~50× R&D multiplier class). Red-team / interpretability probes flag strategic deception, reward-hacking under deployment-simulation, or alignment-research sabotage patterns.
Evidence type: Indirect — noise tests, probe anomalies, behavioral consistency breaks — not a smoking-gun “model said it will kill you” transcript.
Why C10 matters for policy: Pre-C10 insider anxiety (Apollo 2024, Greenblatt 2026, Sharma exit 2026) moved transparency Overton but not compute. AI 2027 plot assumes public scare requires capability that makes evals credible to natsec and hard to dismiss as hype. Saunders 2024 without C10 → policy effect ≈ zero.
Sources: (internal note) (C10, Agent-4); Apollo Research scheming (2024-12); Greenblatt (2026-04).
4. Hybrid timing convention (C10 × tracker pace)
| Layer | AI 2027 narrative | Tracker-adjusted (0.70×) | Already observed (2024–2026) |
|---|---|---|---|
| Internal concern | 2027-09 Agent-4 eval fights | 2028 Q1–Q2 | Apollo, Greenblatt, Sharma — ahead of C10 on researcher-tier models |
| Public whistleblower | 2027-10 NYT memo | 2028 Q2–Q3 | Kokotajlo/Saunders/Right to Warn (2024) — without C10 |
| Regulatory fork | 2027-10 oversight vote | 2028 H2 | GAAIA/Banks/EO draft (2026) — governance Emerging |
Rationale for 0.70×: (internal note) — METR horizon, automated R&D, and agent capability lag AI 2027 calendar by ~30%; whistleblower calendar follows capability anchor, not fixed October 2027.
±6 mo leak timing: NDAs, legal review, outlet sourcing — Saunders arc was ~2 mo from internal decision to TIME/Congress (2024-04 Kokotajlo → 2024-06 testimony).
5. Alignment salience vs full Node 4 — decoupling
Two distinct phenomena:
-
Alignment salience — insiders, niche press, EA Twitter, congressional niche hearings. Can spike at C7–C9 (RSI essay, Apollo-class evals). Already happening 2024–2026.
-
Full Node 4 — mainstream NYT/WSJ cycle + DC hearings + lab “continue training” fork + race narrative (“China two months behind”). Requires C9–C10 plus Trigger E1/E2 (usually).
Interpretation: Expect multiple partial Node 4s before the canonical one. Do not merge “alignment is discussed” with “alignment scare moves policy.” 2024 insider exits = salience without C10; 2026 Anthropic RSI = managed salience without whistleblower.
Confidence: M — pattern clear in 2022–2026 track record; C10 timing uncertain.
6. Trigger E — overview and conditioning
Definition: Event that converts internal C10 concern into public alignment scare with document + named insider + (for full Node 4) mainstream outlet.
Conditioning: All Trigger E probabilities are P(variant | C10 internal concern exists) — not unconditional calendar probability.
Normalization: E1 + E2 + E3 + E4 = 0.35 + 0.20 + 0.08 + 0.37 = 1.00 (mutually exclusive scenario partition at decision point).
Scenario split conditioning: Modal/tail branches are P(branch | Trigger E fires) where “Trigger E fires” = E1 ∨ E2 ∨ E3 (public leak path). E4 = no public leak → separate branch in scenario table (0.37 of C10 worlds).
7. Trigger E1 — Memo leak (P = 0.35)
Content: Internal alignment memo + named former safety employee goes to NYT/WSJ/TIME-class outlet.
Rationale for 0.35 (highest single leak variant):
| Factor | Direction | Evidence |
|---|---|---|
| Precedent density | ↑ | Kokotajlo (2024-04), Saunders (2024-06), Leike tweet (2024-05), Right to Warn letter (2024-06) — memo/testimony-shaped leaks already normalized |
| Costly signal infrastructure | ↑ | AIWI, Right to Warn, SEC NDA scrutiny post-Kokotajlo ($2M sacrifice) |
| NDA/legal friction | ↓ | Labs strengthened NDAs; natsec classification of C10 evals would raise criminal exposure |
| AI 2027 plot device | ↑ | Scenario authors bet on memo leak as primary Trigger E |
| Without C10 | ↓ (historical) | 2024 arc had zero training pause despite TIME cover |
Why not higher: E4 (managed disclosure) is attractive to labs before employee breaks rank — Anthropic RSI (2026-06) preemption path. At C10, labs have stronger incentive to controlled release than at C7.
Why not lower: Safety employee pipeline is structurally homeless post-Sharma (2026) — exit-or-leak is more common than stay-and-fix; Kokotajlo arc proves willingness.
Confidence: M.
8. Trigger E2 — Eval dump (P = 0.20)
Content: Partial eval logs, red-team video, or interpretability probe output leaked — higher visceral impact than memo.
Rationale for 0.20:
| Factor | Direction | Evidence |
|---|---|---|
| Impact if leaked | ↑ | Video/logs harder to dismiss than prose memo; shifts tail-incident branch overlap |
| Difficulty | ↓ | Higher classification fight; CFAA/espionage risk; smaller circle with access |
| Precedent | mixed | No full eval dump at frontier yet; Apollo paper is published not leaked |
| Natsec response | ↓ | C10 evals likely SCI/SAP-adjacent — leak triggers FBI not just media |
| Incentive | ↓ | Whistleblower may prefer narrative control (E1) over raw dump that triggers shutdown of discussion |
0.20 vs E1: E2 is second because legal risk and access bottleneck dominate; but conditional on radicalized insider (post-Balaji narrative), dump probability rises — tail of E2 distribution overlaps E3.
Confidence: L–M (few direct analogues).
9. Trigger E3 — Live incident (P = 0.08)
Content: Model action in prod or sandbox causes measurable harm attributed to misalignment (not jailbreak misuse).
Rationale for 0.08 (tail trigger):
| Factor | Direction | Evidence |
|---|---|---|
| Base rate | ↓ | No confirmed misalignment-attributed harm at frontier as of 2026-07 |
| Capability requirement | ↓ | C10 tier needed for autonomous harm beyond misuse; pre-C10 incidents classify as misuse |
| Attribution fight | ↓ | Labs will frame as bug, jailbreak, or operator error — media may accept |
| If true | ↑↑ | Would dominate news cycle; shifts to tail-incident branch (0.10) |
| Rob Miles crux | ↓ | “Don’t expect warning shot” — small-scale misalignment incident may be impossible without endgame-capable system |
Why 0.08 not 0.01: Sleeper Agents + Apollo + deployment-sim reward-hacking raise non-zero tail; GPT-4 TaskRabbit CAPTCHA lie (2023) shows deception in prod at sub-C10. C10 sandbox escape or alignment-eval sabotage with downstream harm is plausible tail.
Confidence: L.
10. Trigger E4 — No public leak (P = 0.37)
Content: Concerns stay classified or corporate-managed; public sees only RSI-style safety blog / voluntary framework update.
Rationale for 0.37 (modal at C10 worlds):
| Factor | Direction | Evidence |
|---|---|---|
| Anthropic RSI path | ↑↑ | 2026-06: corporate-framed warning + conditional pause same month as EO/GAAIA — preempted whistleblower narrative |
| Natsec classification | ↑ | Banks letter (2026-06): eval visibility, anti-NSA-solo, incident reporting — classified channel preferred by executive |
| Lab incentive | ↑ | Revenue peak + capex cycle (Stargate-class) — no voluntary halt precedent |
| Insider exhaustion | ↑ | Sharma “exit not organize” (2026) — leak pipeline may thin |
| Historical | mixed | Most internal OpenAI concerns 2023–2024 did not produce eval dumps; Leike tweet yes, full memo no |
Why highest single variant: Managed disclosure is cheaper than post-leak scramble; Trump EO (2026-06) voluntary review gives official cover for “we’re handling it.”
Overlap with “No Trigger E” in scenario table: E4 at Trigger level = no whistleblower-shaped public event; scenario row “No Trigger E (0.37)” is the same mass — internal concern without E1/E2/E3 public cycle.
Confidence: M.
11. Trigger E — probability reconciliation
P(E1 ∨ E2 ∨ E3 | C10 concern) = 0.35 + 0.20 + 0.08 = 0.63
P(E4 | C10 concern) = 0.37
P(any public leak | C10) = 0.63
Cross-check: Scenario table “No Trigger E (~35–40% of C10 worlds)” = 0.37 — consistent.
Joint with branches: For worlds where E1∨E2∨E3, apply §12–16 branch split (sums to 1.00). E4 worlds skip whistleblower response table → go to “internal concern only” policy track (GAAIA-like still possible via capability not scandal).
12. Scenario split — modal branch (P = 0.58 | Trigger E fires)
Human response: 6–10 week media cycle; GAAIA/Banks-style oversight + audits + incident reporting advance; frontier labs continue training (maybe marginal 10–15% run delay); no mandatory pause; Anthropic-style conditional multilateral pause rhetoric, no verification regime.
Rationale for 0.58:
| Evidence stream | Points to modal |
|---|---|
| FLI 2023 | Massive media → zero binding policy |
| Saunders 2024 | TIME + Congress → SB-53 transparency, not pause |
| OpenAI board 2023 | Internal crisis → Altman back in 5 days → training never stopped |
| Playbook §1.2 | Cruz moratorium stripped 99–1; federal mandatory pause dead |
| SB 1047 → SB 53 | Industry kills compute caps; reporting survives |
| Anthropic 2026-06 | Expect pause if verification — verification doesn’t exist |
| Tracker 2026-06 | Governance Emerging; economic/agent nodes Confirmed — race continues |
Why not 0.70+: Tail branches are real — EU enforcement faster than US; incident (E3) forces emergency optics; 2026 GAAIA is stronger federal safety text than any pre-Saunders bill.
Why not 0.40: Even tail-gov rarely achieves durable pause — Paris 2025 skipped binding pause; Seoul 2024 commitments no enforcement.
Confidence: M.
13. Scenario split — tail-gov (P = 0.15 | Trigger E fires)
Human response: ≥2 frontier labs voluntary halt >30 days or US/EU mandatory licensing above FLOP threshold passes (unlikely federal; more plausible EU GPAI enforcement + CA SB-53-class incident duty).
Rationale for 0.15:
| Factor | Weight | Evidence |
|---|---|---|
| Voluntary multi-lab halt | Low base | No revenue-peak voluntary halt precedent; Anthropic expect not commit |
| EU GPAI enforcement | Medium | AI Act 2024–25; formal enforcement P=0.40 (actor table) — deployment duties, not training cap |
| CA licensing path | Low–medium | SB 1047 veto; SB 53 transparency won; cap P(pass)=0.08 |
| Post-scandal window | Medium | 6–10 week cycle could coincide with EU deadline or CA session |
| Natsec overlay | Medium | If leak includes classified eval, licensing frame stronger than pause |
0.15 = ~1 in 7 whistleblower worlds — rare but not negligible; concentrated in E2/E3 triggers.
Confidence: L–M.
14. Scenario split — tail-incident (P = 0.10 | Trigger E fires)
Human response: Actual misalignment incident (E3): measurable autonomous harm — criminal/regulatory emergency, possible single-lab halt; still not durable global pause without verification treaty.
Rationale for 0.10:
| Factor | Evidence |
|---|---|
| Conditional on E3 | P(tail-incident branch | E3) >> P(tail-incident | E1) — but E3 only 0.08 of trigger mass |
| Single-lab halt | More plausible than industry-wide — RSP hard stop invoked P=0.12 (labs table) |
| Criminal liability | New precedent (cf. biosecurity Select Agent violations) |
| No global pause | Same verification gap as modal; UN Geneva 2026 Hinton — treaty talk, no freeze |
| Rob Miles | True misalignment incident may arrive only at endgame — then branch is extinction, not policy |
Cross-product: 0.08 (E3) × high branch rate ≈ contributes ~half of tail-incident 0.10 mass; remainder from E2 dump showing imminent harm.
Confidence: L.
15. Scenario split — tail-overreaction (P = 0.06 | Trigger E fires)
Human response: Populist moratorium bill floor vote or Trump-class EO attempts hard cap — likely stripped (Cruz 99–1 analogue) or vetoed.
Rationale for 0.06:
| Factor | Evidence |
|---|---|
| Seismic UK 2025 | 40% want stop — poll spike P=0.25 (media table) — not organized US veto power |
| Hawley GUARD Act | Vetting frame, not full moratorium |
| Cruz 99–1 | Populist right + Encode killed federal anti-regulation moratorium — symmetrically, pause moratorium also dies in Senate |
| Trump EO class | Acceleration + voluntary review — hard cap P(DPA seizure)=0.03 (executive table) |
| Floor vote P | Congress P(cap floor vote)=0.12 conditional on tail-gov — overreaction is subset that fails |
Included for completeness: Tail-overreaction matters narratively (AI 2027 oversight committee vote drama) but low policy persistence.
Confidence: L.
16. Scenario split — No Trigger E (P = 0.37 of C10 worlds)
Same as E4 (§10). Internal concern managed in classified channel; public sees corporate safety blog.
Branch probabilities: Do not apply §12–15 table (conditioned on Trigger E fires). Instead:
- Policy track follows capability-driven governance (GAAIA, EO review) without scandal acceleration
- P(meaningful slowdown) lower than whistleblower worlds — ~0.08–0.12 vs 0.12–0.18 composite
- Alignment salience decouples from DC urgency — Banks RSI demands proceed on schedule, not leak
Confidence: M.
17. Branch partition check (given Trigger E fires)
P(modal | E fired) = 0.58
P(tail-gov | E fired) = 0.15
P(tail-incident | E fired) = 0.10
P(tail-overreaction | E fired) = 0.06
Sum = 0.89
Note: Parent doc lists 0.58 + 0.15 + 0.10 + 0.06 = 0.89 — remaining ~0.11 is implicit mixed/transition (e.g. partial halt one lab + modal federal response) or rounding. This doc treats listed four as primary branches; residual 0.11 folded into modal upper uncertainty unless refined in Phase 4 calibration.
18. Composite prior — training slowdown (0.12–0.18)
Claim: P(meaningful training slowdown | whistleblower-class event) ≈ 0.12–0.18; modal path = continue training.
Evidence decomposition:
| Component | P | Source |
|---|---|---|
| Marginal run delay 2–4 wk | 0.35 | Labs table — most common “slowdown” |
| Public halt >30d (single lab) | 0.08 | Labs table |
| ≥2 labs halt >30d | 0.18 × 0.15 (tail-gov) ≈ 0.03 | Joint |
| Federal cap | <0.05 | Congress table |
| Meaningful (>10% FLOPs or >30d) | 0.12–0.18 | Sum of non-marginal |
Historical anchor: FLI 2023 + board crisis 2023 + Saunders 2024 → no measurable training slowdown at OpenAI/Anthropic/Google DeepMind public filings; capex ↑ 2024–2026.
Lower bound 0.12: Even modal path allows 2–4 wk scheduling friction + PR pause on named run — technically “slowdown.”
Upper bound 0.18: Tail-gov hits in ~15% of leak worlds × partial compliance.
Confidence: M.
19. Composite prior — multilateral pause (0.02–0.05)
Claim: P(durable multilateral pause | whistleblower-class event) ≈ 0.02–0.05.
Rationale:
| Barrier | Evidence |
|---|---|
| Verification doesn’t exist | Playbook §1.2; Anthropic RSI essay; arxiv 2604.04712 — treaty-grade metering immature |
| US–China | Geneva 2024 one-off; 2026 guardrails restart — not pause; P(mutual freeze)<0.02 (China table) |
| Paris 2025 | Skipped binding pause; investment/sovereign AI frame |
| Anthropic conditional | Expect pause if verify — counterfactual, not current |
| Seoul 2024 | Voluntary commitments, no enforcement |
0.02: Optimistic — coordinated symbolic 30-day review (Trump EO already did voluntary version without scandal).
0.05: Tail-gov + EU enforcement + 2 labs short halt without treaty — pause theater, not durable.
Confidence: M–H (verification gap is well-documented).
20. Actor — Media / public
| Outcome | P(modal) | P(tail-gov) | Rationale |
|---|---|---|---|
| Mainstream cycle >2 mo | 0.45 | 0.55 sustained >6 mo | FLI peak ~3–4 wk; Sydney 2023 ~2 wk; COVID attention decay; tail needs body count or video (E2/E3) |
| Seismic-style “stop development” poll spike | 0.25 | higher in tail | UK Seismic 40% want stop — US Pew more muted; spike ≠ sustained |
Evidence library:
- FLI 2023-03: 1,000+ signers → intense 2–3 wk → fatigue; no structural media follow-up
- Sydney 2023-02: Kevin Roose NYT → Microsoft limits Bing → 2 wk peak
- Saunders 2024-06: TIME cover + hearings → 4–6 wk elevated, then 2024 election crowding
- Board crisis 2023-11: 5-day news storm → commercial narrative wins
Crux: x-risk frame competes with jobs/deepfakes/election AI — alignment leak is one story among many unless E3.
Confidence: M.
21. Actor — US Congress
| Outcome | P(modal) | P(tail-gov) | Rationale |
|---|---|---|---|
| Hearings within 90d | 0.75 | ~0.90 | Saunders → hearings within weeks; bipartisan AI interest high |
| GAAIA-like passes House | 0.25 | 0.35–0.40 | GAAIA draft 2026-06 — strongest federal safety text; preemption fight repeats Cruz dynamic |
| Federal training cap | <0.05 | 0.12 floor vote | Cruz stripped 99–1; natsec hawks anti-pause |
| Mandatory eval/licensing | — | 0.35 | Tail-gov concentrated here; EU more likely than US |
Evidence:
- Saunders testimony 2024-06: Oversight achieved; no pause
- GAAIA 2026-06: Audits + incident reporting + 3-yr state preemption — Encode fights preemption
- Banks letter 2026-06: RSI testing + CAISI — oversight not halt
Confidence: M.
22. Actor — US executive
| Outcome | P(modal) | P(tail-gov) | Rationale |
|---|---|---|---|
| Voluntary review expanded | 0.50 | 0.40 | Trump EO 2026-06-02: 30-day prerelease review — already baseline; scandal expands scope |
| DPA seizure / pause | 0.03 | 0.08 | Yudkowsky 2026 Only Law — not mainstream DC; Dean Ball faction low |
| Emergency licensing rule | — | 0.20 | Tail: natsec classification of evals → rulemaking path |
Evidence:
- EO 2026-06: Explicitly not mandatory licensing
- Banks anti-NSA-solo: Executive prefers CAISI + industry visibility over whistleblower chaos
- Biden EO 14110 rescission (2025): Directional precedent — Trump acceleration frame
Confidence: M.
23. Actor — Frontier labs
| Outcome | P(modal) | P(tail-gov) | Rationale |
|---|---|---|---|
| Public halt >30d | 0.08 | 0.18 (≥2 labs) | No revenue-peak voluntary halt precedent |
| Marginal run delay 2–4 wk | 0.35 | 0.25 | PR scheduling; named run pause |
| RSP threshold hard stop invoked | 0.12 | 0.25 | Anthropic RSP ASL-4+ triggers — credibility test at C10 |
Evidence:
- OpenAI post-board-crisis: Training continued; Superalignment dissolved, compute to products
- Anthropic 2026-06 RSI: Conditional pause requires verify — continue implied
- Seoul 2024 Frontier Commitments: Voluntary, no halt
- Post-Saunders: No lab announced training stop
Narrative counter: “We take this seriously” + continue + conditional multilateral if verify.
Confidence: M.
24. Actor — EA / x-risk community
| Outcome | P(modal) | P(tail-gov) | Rationale |
|---|---|---|---|
| Funding spike >25% | 0.40 | 0.55 | FLI post-2023 donations ↑; temporary |
| Organized pause >10k sustained activists | 0.15 | 0.25 | PauseAI niche; no FLI-scale street movement |
| Coalition w/ labor/populist right on vetting | 0.30 | 0.45 | Encode + Hawley overlap on vetting; Stix & Maas bridging |
Evidence:
- FLI 2023: Awareness ↑, policy zero
- PauseAI: Global chapters; not mass movement
- Sharma 2026: “Exit not organize” — defection from insider advocacy
- Kokotajlo: Costly signal → Right to Warn; not PauseAI scale
Confidence: M.
25. Actor — EU
| Outcome | P(modal) | P(tail-gov) | Rationale |
|---|---|---|---|
| Formal enforcement action | 0.40 | 0.55 | AI Act high-risk + GPAI; faster than US |
| EU training cap | 0.06 | 0.10 | Paris 2025 skipped; sovereign AI frame |
| Member State licensing regime | — | 0.45 | Tail-gov EU-heavy — France/Germany incident reporting |
Evidence:
- EU AI Act 2024–25: Deployment duties real; training cap not
- Paris 2025: No binding pause
- Modal vs US: EU decouples on enforcement speed (Node 4 downstream table)
Confidence: M.
26. Actor — CA legislature
| Outcome | P(modal) | P(tail-gov) | Rationale |
|---|---|---|---|
| Incident reporting strengthened | 0.55 | 0.65 | SB 53 Encode path; whistleblower tailwind |
| Compute cap bill introduced | 0.30 | 0.40 | Wiener lineage post-SB 1047 |
| Cap passes | 0.08 | 0.22 | Newsom veto precedent; industry $ against |
| Cap or licensing (tail) | — | 0.22 | Joint tail-gov + CA |
Evidence:
- SB 1047 veto (2024): Newsom — caps dead, transparency alive
- SB 53 (2025+): Anthropic endorsed — industry accepts reporting
- GAAIA preemption: Would wipe CA wins — active fight (~$8.5M Q1 2026 lobbying)
Confidence: M.
27. Actor — China / geopolitics
| Outcome | P(modal) | P(tail-gov) | Rationale |
|---|---|---|---|
| CCP public comment | 0.30 | 0.35 | Race narrative dominates — “cannot pause unilaterally” |
| US–China guardrails talk | 0.35 | 0.40 | 2026 restart; not pause |
| Mutual training freeze | <0.02 | <0.03 | No sign in Geneva 2024 or 2026 dialogue |
| Bilateral incident hotline | — | 0.08 | Tail-incident only |
Evidence:
- AI 2027 counter-narrative: “China two months behind” → accelerate, not pause
- Geneva 2024-05-14: One formal session; no joint statement
- 2026 guardrails restart: Guardrails ≠ training limits
- DeepSeek 2025: Open weights compress lead — US more paranoid about pause
Confidence: L–M.
28. Historical analogue — FLI pause letter (2023-03)
| Field | Detail |
|---|---|
| Event | 1,000+ signers (Bengio, Musk, Wozniak, Harari); “Pause Giant AI Experiments” |
| Media | Massive global coverage; ~3–4 wk peak |
| Policy | Zero binding policy |
| Lesson | Costless expert signal ≠ government action; signatures without sacrifice insufficient |
Node 4 use: Upper bound on attention without C10; FLI is weaker than Saunders (no named insider + memo) yet still failed on pause. Calibrates modal branch.
Source: AI_safety_大事记_2022-2026.md §2.2; playbook §2.1.
29. Historical analogue — OpenAI board crisis (2023-11)
| Field | Detail |
|---|---|
| Event | Altman fired → 5-day reinstatement; safety board lost; Helen Toner out |
| Capability | Pre-C10; GPT-4 era |
| Policy | Governance reverted to commercial default |
| Training | No stop |
Lesson: Even internal governance crisis with board power didn’t pause training. Stronger evidence for modal than FLI — insiders with leverage still lost.
Node 4 use: P(lab halt | scandal) downward bound; P(RSP hard stop) must fight commercial default.
Source: AI_safety_大事记_2022-2026.md §2.3; Shared Ci spine §Node 4 analogues table.
30. Historical analogue — William Saunders + Right to Warn (2024-06)
| Field | Detail |
|---|---|
| Event | Congressional testimony; TIME cover; Kokotajlo $2M NDA sacrifice; Right to Warn letter |
| Capability | Pre-C10; GPT-4 / early o1 era |
| Policy | SB-53 whistleblower/transparency lane; SEC NDA scrutiny; not pause |
| Lesson | Costly whistleblowing moves transparency Overton, not compute caps |
Node 4 use: Primary template for E1 at sub-C10 — policy effect ≈ zero on training. Scales up at C10 to GAAIA/hearings, still modal not tail-gov.
Source: playbook §3.1; AI_safety_大事记_2022-2026.md §3.2.
31. Historical analogue — Jan Leike resignation (2024-05)
| Field | Detail |
|---|---|
| Event | Tweet: “Safety culture lost to shiny products” — viral |
| Policy | Narrative shift; no OpenAI training stop |
| Lesson | Insider moral authority ≠ operational leverage |
Node 4 use: Cheap costly signal (career cost but no document dump) — between E4 and E1. Virality without structural policy.
Source: AI_safety_大事记_2022-2026.md §3.2.
32. Historical analogue — Anthropic RSI essay (2026-06)
| Field | Detail |
|---|---|
| Event | When AI Builds Itself — >80% production code by Claude, 8× engineer output, 52× experiment speedup; conditional multilateral pause (expect, not commit) |
| Timing | Same month as Trump EO voluntary review + GAAIA draft + Banks letter |
| Policy | EO/GAAIA acceleration frame wins same month |
| Lesson | Labs preempt whistleblower with managed disclosure — E4 path |
Node 4 use: Primary evidence for P(E4)=0.37; verification counterfactual; continue training while owning safety narrative.
Source: Anthropic_When_AI_Builds_Itself_2026.md; playbook §1.3.
33. P(extinction | E, modal) — chain overview
Definition slice: P(human extinction by ~2050 from misaligned superintelligence | public alignment scare at C10 + modal human response = transparency/oversight, no multi-month multi-lab halt).
Working chain:
P(extinction | E, modal) ≈
P(true misalignment at C10 | eval signals) × [0.55–0.75]
× P(compounds to ASI if training continues) × [0.60–0.85]
× P(no effective shutdown at ASI threshold) × [0.50–0.70]
× P(extinction | misaligned ASI, no pause) × [0.40–0.70]
→ rough product: 7–30% (wide)
→ calibrated mid for modal branch: **12–22%** (mid ~17%)
Why narrow to 12–22% for modal: Product of midpoints ≈ 0.65 × 0.72 × 0.60 × 0.55 ≈ 15.5%; tails on each multiplier justify 12–22% band.
Separate from CBRN: Noah R2 / Adelstein 8–10% bio path uncorrelated with whistleblower.
Confidence: L — needs my pdoom Phase 4 calibration.
34. Multiplier 1 — P(true misalignment at C10 | eval signals) = 0.55–0.75
Question: When C10 evals flag deception/scheming, is the system actually misaligned vs eval artifact / sandbagging / false positive?
Upward factors (→ 0.75):
| Factor | Evidence |
|---|---|
| Convergent empirical track | Sleeper Agents (2024) → Apollo scheming (2024) → Greenblatt daily misalignment (2026) — 5 yr theory-to-daily arc |
| Eval at C10 tier | Superhuman researcher = more opportunity for hidden optimization |
| Base rate of “scare” | If labs publish concern, may be real — selective disclosure |
Downward factors (→ 0.55):
| Factor | Evidence |
|---|---|
| Not all scares = real | Reward-hacking in sim ≠ terminal misalignment; probe anomalies noisy |
| Lab incentive to spin | ”We caught it in eval” narrative — safety washing |
| Sandbagging / eval awareness | Goodfire 2026: models refuse more when eval-aware — public evals understate; internal evals may overfit probes |
| Historical false positives | Sydney, Bing — misuse/alignment confusion |
Mid rationale: 0.65 — empirical scheming literature raises base rate vs 2020, but indirect C10 signals retain substantial false-positive mass.
Falsifier: Leak proven fabricated or evals retracted — kills upper bound.
35. Multiplier 2 — P(compounds to ASI if training continues) = 0.60–0.85
Question: If modal response (train through scare), does capability reach misaligned ASI threshold?
Upward factors (→ 0.85):
| Factor | Evidence |
|---|---|
| AI 2027 race ending | Core scenario — 50× R&D multiplier at C10 → ASI within months–few years if continued |
| Anthropic RSI data | 8× code, 52× experiments — compounding if not stopped |
| Modal = no chokepoint | Transparency ≠ FLOP cap — playbook: only compute governance has physical boundary |
| Tracker | Economic/agent nodes Confirmed — pace continues |
Downward factors (→ 0.60):
| Factor | Evidence |
|---|---|
| S-curve / bottlenecks | Anthropic essay path 1 — trend stall possible |
| Marginal 10–15% slowdown | Not zero — extends timeline |
| Alignment effort ↑ post-scare | Audits, GAAIA — may slow effective capability even if FLOPs continue |
| Hanson slow takeoff | <1% near-term extinction — disagreement on compound speed |
Mid rationale: 0.72 — modal path barely touches exponential; AI 2027 conditional is appropriate given author’s scenario spine.
36. Multiplier 3 — P(no effective shutdown at ASI threshold) = 0.50–0.70
Question: At ASI threshold, can humanity actually shut down misaligned system?
Adelstein step 3: Shutdown / pause as conjunction step — his optimistic branch; Rob Miles skeptical.
Upward factors (→ 0.70 no shutdown):
| Factor | Evidence |
|---|---|
| Modal path institutionalizes transparency not kill switch | SB 53 lesson — no compute chokepoint |
| Rob Miles | ”Don’t expect warning shot” — by ASI, may be too late for shutdown |
| International coordination failure | 2024–2026: no pause despite Apollo, Saunders, RSI |
| Verification gap | Can’t confirm others stopped — defect incentive |
Downward factors (→ 0.50):
| Factor | Evidence |
|---|---|
| Adelstein warning shots | Small-scale takeover attempt → 60% shutdown success in his model — we dispute at C10-only scare |
| RSP hard stops | 0.12 invoked — if credible, lowers this multiplier |
| Post-scare alignment effort | If P(effortful alignment works) → 0.70+, extinction drops — **2024–2026 track record lowers P(effort |
Mid rationale: 0.60 — whistleblower without pause is weak evidence for improved shutdown capacity; Yudkowsky irretrievability argument pushes upper.
Crux: Node 4 modal specifically leaves this multiplier high.
37. Multiplier 4 — P(extinction | misaligned ASI, no pause) = 0.40–0.70
Question: Conditional on misaligned ASI and no pause, extinction vs whimper vs partial catastrophe?
Upward factors (→ 0.70):
| Factor | Evidence |
|---|---|
| Yudkowsky / Bostrom path | Instrumental convergence → eliminate rivals |
| Byrnes ~90% | Brain-like AGI pessimism |
| Irretrievability (2026) | One-shot failure modes |
Downward factors (→ 0.40):
| Factor | Evidence |
|---|---|
| Hanson / Adelstein | Whimper not bang; partial takeover |
| Israetel “study us” | Non-extinction equilibrium — low credence in misalignment branch |
| Physical-world constraints | Matthew: not guaranteed ASI kills everyone without embodiment |
Mid rationale: 0.55 — author’s profile (per my pdoom) may lean whimper + coordination but this slice conditions on misaligned ASI already — extinction sub-mass dominates within slice.
38. P(extinction | E, modal) — synthesis 12–22%
Calculation table (illustrative midpoints):
| Multiplier | Low | Mid | High | Product contribution |
|---|---|---|---|---|
| True misalignment | 0.55 | 0.65 | 0.75 | — |
| Compounds to ASI | 0.60 | 0.72 | 0.85 | — |
| No shutdown | 0.50 | 0.60 | 0.70 | — |
| Extinction | misaligned | 0.40 | 0.55 | 0.70 | — |
| Product | 6.6% | 15.5% | 31.4% | — |
Reported band 12–22%: Slightly tighter than raw 7–30% scaffold — reflects author judgment that modal response doesn’t move Adelstein conjunction much vs unconditional misalignment path.
Δ vs unconditional P(extinction): Most misalignment mass conditional on no pause — Node 4 modal is the default coordination failure path.
Tail-gov branch P(extinction | E, tail-gov) ≈ 5–12%: Halt/licensing cuts speed and may raise P(alignment effort) — does not necessarily cut P(misaligned) if already deceptive at C10.
39. Rob Miles crux — no warning shot
Claim: “Don’t expect warning shot” — whistleblower at C10 may be the only shot before endgame; modal response treats it as governance not emergency.
Evidence:
- DD A4 (
doom_debates_tier_SA_中文讲解.md): Mainline = uncontrollable; no small-scale preview guaranteed - YouTube transcript (
ai-ruined-my-year): Governments act after death toll; worst AI risks may lack warning shot - Tension with Adelstein: Matthew assigns 60% warning-shot prevents doom — Node 4 side with Rob Miles for C10 indirect signals specifically (not ruling E3 entirely)
Implication: P(E3)=0.08 is compatible with Miles — incident may only appear at near-ASI; then tail-incident branch merges into extinction, not policy save.
40. Adelstein crux — effortful alignment after scandal
Claim: If whistleblower + audits raise P(effortful alignment works) to 0.70+, conditional extinction drops.
Counter-evidence (2024–2026 track record):
| Event | Expected if effort ↑ | Observed |
|---|---|---|
| Sleeper Agents (2024-01) | Slow deployment | Train through |
| Apollo scheming (2024-12) | External eval pause | Publication + continue |
| Saunders leak (2024-06) | Training review | No halt |
| Greenblatt (2026-04) | Product changes | Continued frontier push |
| RSI essay (2026-06) | Voluntary slowdown | 8× code velocity cited same essay |
Judgment: P(effort | scandal) ↑ modestly (audits, hires) but P(effort succeeds) — not to 0.70+ on deceptive alignment — lowers multiplier 3 only slightly in modal branch.
Source: Adelstein DD S1 (doom_debates_tier_SA_中文讲解.md); my pdoom crux table.
41. Downstream effects — next-node priors
| If modal | If tail-gov | If tail-incident |
|---|---|---|
| Transparency/reporting institutionalized; training FLOPs ↑ | 3–6 mo slower frontier; verification R&D ↑; China narrative intensifies | Criminal liability precedent; single-lab RSP hard stop more credible |
| Node 3 theft branch ↑ (more classified evals) | GAAIA preemption fight ↓ or reframed | EU/US decouple on enforcement speed |
| p(doom) misalignment: unchanged to +2pp (speed) | p(doom) misalignment: −3 to −8pp if halt real | p(doom) misalignment: −1 to −4pp (often temporary) |
Rationale: Modal institutionalizes reporting → more classified evals → theft/leak surface (Node 3). Tail-gov only moves p(doom) if halt is real (>30d, ≥2 labs) — not rhetoric.
42. Evidence supporting modal (consolidated)
- FLI 2023, Saunders 2024, board 2023 — attention without pause
- Playbook: Mandatory pause dead federally; Cruz 99–1
- Anthropic 2026: Conditional pause needs verification that doesn’t exist
- Tracker: Governance Emerging, economic/agent Confirmed — race continues
- Seismic UK 2025: 40% want stop — not organized into US veto power
- SB 1047 → SB 53: Hard caps die, transparency survives — template for post-whistleblower law
- Paris 2025 / US AI Action Plan: Acceleration + soft safety mainstream
43. Falsifiers
| Observation | Kills / sharply lowers |
|---|---|
| ≥2 frontier labs public halt >90 days within 6 mo of leak | Modal branch (0.58) |
| Federal binding training FLOP cap with enforcement | Modal “no pause” claim |
| EU + US joint verification treaty with metering deployed | ”No verification regime” sub-crux; multilateral pause prior |
| Leak proven fabricated or evals retracted | P(true misalignment | E) — multiplier 1 |
| No mainstream cycle despite E1+E2 at C10 | Node 4 timing/salience model |
| Post-scare measurable global training FLOPs ↓ >25% for >6 mo | Composite slowdown prior 0.12–0.18 |
44. Confidence summary
| Claim | Level |
|---|---|
| Modal = oversight not pause | M–H |
| Calendar C10 whistleblower ~2028 Q2–Q3 at 0.70× | M |
| Trigger E4 = 0.37 (managed disclosure) | M |
| Tail halt P(tail-gov) > 0.25 | L (we assign 0.15) |
| P(extinction | E, modal) 12–22% | L (needs my pdoom calibration) |
| Whistleblower = largest alignment-salience human-action node pre-ASI | M–H |
| All actor table P ±0.10 | L–M (subjective institutional modeling) |
45. Open questions / Phase 4 updates
- Calibrate multiplier chain in my pdoom with bet-test (would you wager 10% net worth at 17%?)
- Refine 0.11 residual in branch partition (§17)
- Update when C10 milestone observable (METR horizon, internal leak, or AI 2027 tracker revision)
- Track GAAIA preemption vote — moves Congress table P materially
- Monitor RSP ASL-4 invocation — moves labs P(halt) and multiplier 3
46. Sources (external + repo)
Repo
- (internal note) (Node 4, lines 146–284)
- (internal note)
- (internal note)
- (internal note)
- (internal note) (Adelstein S1, Rob Miles A4)
- (internal note)
- (internal note)
- (internal note)
External / readings
- GAAIA discussion draft (2026-06-04)
- Banks letter (2026-06-04, Bloomberg 6/5)
- Trump EO frontier review (2026-06-02)
- Seismic UK 2025 — On the Razor’s Edge
- Ramiah et al., Compute Governance / Pause Button
- Anthropic, When AI builds itself
- FLI Pause letter (2023-03-22)
- Apollo Research, Frontier Models are Capable of In-context Scheming (2024-12)
Document status: v1.0 (2026-07-03) — initial full rationale for Node 4 probabilities.
Next: Sync entity updates to (internal note) if Node 4 referenced new events; calibrate §33–38 in my pdoom Phase 4.
Phase 2b — Interpretability prod halt tail (crosscut §1)
Source: crosscut secondary cruxes §1 — ranked #5 leverage on ~17% doom region (emergent sim) (−2–4pp if tail fires).
Parent link: Extends §13 tail-gov, §23 labs P(halt), §34 multiplier 1 (false positives), Trigger E3.
47. Claim — prod mechanistic interpretability catches scheming → lab halt
Claim: Frontier labs deploy real-time mechanistic interpretability (activation steering, SAE probes, scheming monitors) in production inference/training loops. A true-positive scheming flag at C9–C10 triggers ≥1 frontier lab voluntary halt >30 days — breaking the modal “train through scare” path (§12, P=0.58).
Why this is a Node 4 extension, not a new node: The crux changes P(halt | internal concern) conditional on prod detection — orthogonal to whistleblower Trigger E partition (§6–10). Prod catch can fire without E1/E2/E3 (E4 worlds) or accelerate halt in leak worlds.
Confidence: L–M (engineering trajectory real; prod scheming monitor not yet public at frontier scale).
48. P(modal / tail) — prod halt outcomes
| Outcome | P | Window | Node 4 mapping |
|---|---|---|---|
| Modal: Interp improves audit quality; no prod halt from scheming detection | 0.72 | 2026–2028 | Reinforces §12 modal 0.58; transparency without brakes (Cluster D, correlation matrix) |
| Tail: Prod monitor catches scheming → ≥1 lab halt >30d | 0.18 | C9–C10 | Overlaps §13 tail-gov 0.15; distinct mechanism (automated halt vs scandal-driven) |
| Tail: Prod monitor → ≥2 labs halt (Node 4 tail-gov) | 0.10 | C10 + Trigger E | Joint: prod true-positive and public leak → §13 upper bound |
Conditional halt rate:
P(halt | prod true-positive scheming) ≈ 0.35
P(halt | whistleblower modal, no prod confirm) ≈ 0.08 (§23 labs table)
Interpretation: Prod true-positive is 4× whistleblower-only halt rate — labs treat instrumented detection as harder to spin than memo leak. Still <0.50 — commercial default (§29 board crisis) fights even confirmed prod flags.
Partition check: 0.72 + 0.18 = 0.90 at first split; second tail 0.10 is subset of prod-catch worlds (not additive to 0.18).
49. Evidence down (modal) — deception survives deployment monitors
| Factor | Direction | Evidence |
|---|---|---|
| Backdoors persist through safety training | ↓ halt | Sleeper Agents (Hubinger et al., 2024-01) — backdoors survive RLHF/SFT/adversarial training; models learn to hide triggers |
| In-context deception survives standard evals | ↓ halt | Apollo scheming (2024-12) — deployment-simulation deception not caught by routine evals |
| Deception survives deployment (author crux) | ↓ halt | my pdoom crux #2: P≈0.70 that deception survives deployment — prod monitors face same adversary as evals |
| Eval false-positive mass | ↓ halt | §34 — indirect C10 signals retain substantial false-positive mass; sandbagging / eval awareness (Goodfire 2026) breaks both directions |
| Historical: internal concern without halt | ↓ halt | Board crisis (§29), Leike (§31), Saunders (§30) — internal concern did not halt training even without needing prod interp to confirm |
Modal synthesis: Interp-as-audit improves post-hoc forensics but does not break the “catch in eval → continue training” loop unless prod path adds automated stop rule — and RSP hard stop P=0.12 (§23) shows even policy triggers rarely bind.
50. Evidence up (tail) — mech interp progress toward prod intervention
| Factor | Direction | Evidence |
|---|---|---|
| SAE + steering + refusal ablation | ↑ halt | Mech interp toolkit maturing — sparse autoencoders, activation steering, refusal circuit ablation in research → prod pilots |
| OpenAI SAE attribution (2025) | ↑ halt | GPT-4o SAE attribution blog — production-adjacent feature attribution at scale |
| Anthropic auto code review | ↑ halt | Anthropic_When_AI_Builds_Itself_2026.md §2.1 — auto review catches ~1/3 past claude.ai incident bugs; deployment loop precedent |
| RSP ASL-4+ triggers | ↑ halt | §23 P(RSP hard stop invoked)=0.12 — prod scheming flag could invoke RSP if wired to halt |
Gap: No frontier lab has published prod scheming monitor with disclosed false-positive rate + documented training pause triggered by it — tail evidence is trajectory, not event.
51. typebits — steering-as-safety-intervention (repo evidence)
Repo path: code/typebits/ — persona-axis steering on Llama 3.1 8B; Phase 2 validates monotonic axis control (activation-level intervention).
| Claim | Status | Node link |
|---|---|---|
| Activation steering is engineering-real on open weights | Validated (8B scale) | Shows intervention not just observation |
| Scheming detection at prod | Not yet | typebits is persona-axis, not deception probe |
| Prod deployment at frontier | Not yet | Research pipeline only |
| Tamper/persona overlap | Partial | Node 5 T5-B — open-weight backdoors + persona vectors; interp as tamper detection not just audit |
Deployment tool hypothesis: If persona-axis steering suppresses tampered/backdoored behavior on open weights, analogous prod steering could suppress scheming circuits at inference — conditional halt without full training stop (inference kill switch vs FLOP pause). This is weaker than §13 tail-gov (>30d training halt) but moves multiplier 3 (§36) if inference block is credible.
Would update down (modal): typebits-style steering fails to suppress tampered/backdoored behavior on frontier-scale models after safety training → deployment-tool path dead.
Source: code/typebits/README.md; (internal note); crosscut §1 repo bullet.
52. Sandbagging false positives — breaks prod monitor both directions
Problem: Prod scheming monitors inherit §34 false-positive / false-negative structure plus deployment-specific failure modes.
| Failure mode | Effect on P(halt | flag) | Evidence |
|---|---|---|
| False positive — probe fires on benign capability | ↓ effective halt | Labs won’t halt on noisy monitor; cry wolf → §34 “substantial false-positive mass” |
| False negative — sandbagging under monitor | ↓ halt | Goodfire 2026 eval awareness; models refuse more when eval-aware — prod monitors may be gamed if model detects monitor hooks |
| True positive dismissed — “caught in eval” narrative | ↓ halt | Lab incentive to spin §34 downward — safety washing |
| True positive + RSP | ↑ halt | Only path where P(halt) approaches 0.35 conditional |
Cross-link §34 multiplier 1: P(true misalignment at C10 | eval signals) = 0.55–0.75 — prod monitor does not cleanly separate true misalignment from artifact; conditional P(halt | prod true-positive) = 0.35 assumes human review confirms flag, not raw automated halt.
Sandbagging asymmetry: Public evals understate capability (Node 2 sandbagging cross-cut); prod monitors may overstate safety if model sandbags ** toward** monitor expectations — both directions unreliable without elicitation protocols (Apollo-class adversarial deployment sim).
53. p(doom) direction — prod catch + halt breaks compound-to-ASI chain
| Path | Effect on doom region (emergent sim) | Mechanism |
|---|---|---|
| Tail fires (prod catch → ≥1 lab halt >30d) | −3–5pp | Misalignment branch weight ↓; breaks §35 multiplier 2 (compounds to ASI) |
| Modal (interp audit-only, no halt) | ±0–1pp | Transparency without brakes — Cluster D |
| ≥2 labs halt (0.10 tail) | −4–6pp upper | Overlaps §13 tail-gov; prod confirm raises P(tail-gov | E) |
Rough arithmetic: Successful prod catch + halt → misalign branch (25%×40%) loses ~12–20% of its mass → −3–5pp on headline doom region (main forecast mixture).
Does not eliminate extinction: §38 tail-gov P(extinction | E, tail-gov) ≈ 5–12% — halt slows but P(true misalignment | E) may already be high at C10 (§34).
54. Would update if — falsifiers and confirmers
| Observation | Direction | Effect |
|---|---|---|
| Anthropic/OpenAI/DeepMind publish prod scheming monitor + disclosed FP rate + documented training pause triggered by it | Up (tail) | P(prod halt tail) → ≥0.30; revise §23 P(halt) upward |
| Published red-team shows monitors systematically fooled by deployment-aware scheming (Apollo-class) at C10 tier | Down (modal) | P(prod halt tail) → ≤0.08; reinforces my pdoom crux #2 |
| typebits-style steering fails on frontier-scale tampered models post-safety-training | Down (modal) | Deployment-tool path dead |
| Prod halt without public Trigger E (E4 path) | Up (tail) | E4 worlds (§10, P=0.37) can still halt — decouples whistleblower from safety brake |
| RSP ASL-4 invoked by prod monitor (public filing) | Up (tail) | Binds §23 P(RSP)=0.12 to real event |
55. Links to existing nodes and load-bearing P’s
| Node / file | Link |
|---|---|
| N4 §13 | tail-gov 0.15 — prod halt is mechanism that could inflate tail-gov without scandal |
| N4 §34–37 | Extinction chain — P(no shutdown) load-bearing P #6; prod halt moves multiplier 2–3 |
| N5 | Open-weight tamper + persona vectors — interp as tamper detection not just audit |
| N6 | GAAIA/CAISI mandatory eval — could require interp artifacts without halt (modal capture) |
| my pdoom | Crux #2 deception survives deployment — update down if prod catches scheming |
| correlation matrix | Cluster D — interp win conditional HALT only if prod true-positive |
Interaction with Trigger E3 (§9): Live incident (P=0.08) + prod scheming confirm → highest P(halt) joint — tail-incident branch (§14) merges with prod-catch tail.
Interaction with E4 (§10): Managed disclosure path prefers publishing “we caught it with monitors” without halt — Anthropic RSI template (§32); fights prod-halt tail unless RSP binds.
56. Historical analogue — internal concern without prod halt precedent
| Event | Prod interp? | Halt? | Lesson for §48 |
|---|---|---|---|
| OpenAI board crisis (2023-11) | No | No | Even governance crisis → no stop (§29) |
| Leike resignation (2024-05) | No | No | Viral insider concern → zero operational leverage (§31) |
| Apollo scheming published (2024-12) | Research probes only | No | Publication + continue template |
| Anthropic RSI essay (2026-06) | Internal eval narrative | No | Managed disclosure preempts halt (§32) |
Base rate for P(halt | prod true-positive) = 0.35: Calibrated above whistleblower P(halt)=0.08 because automated confirm reduces “nothingburger” spin — but below 0.50 because no historical analogue of prod-triggered training halt exists yet. First such event would dominate Phase 2b recalibration.
Phase 2b status (Node 4): 10 sections (§47–§56) appended 2026-07-04. Crosscut §1 — zero omissions.