I’ve felt torn for a while. In Silicon Valley I talk about AI every day and read material that puts capability jumps around 2027. Outside the Bay and a few Chinese cities, most of the world still hasn’t really been touched by AI.
The stakes are huge. People say utopia or dystopia. Preparation has not kept up—not just technically, mostly institutionally. E.O. Wilson’s line still fits: Paleolithic emotions, medieval institutions, god-like technology. AI widens the mismatch.
Everyone in the Bay talks about AI; most people stay inside their interest radius—investments, layoffs, jobs. Few seriously think about AI’s impact on society. That disconnect bothers me.
This essay is about which problems we need to solve—from long-run x-risk to near-term socioeconomic impact. How critical is each, what’s stuck, what’s the clock, what do success / half-success / failure look like?
Ajeya Cotra put it bluntly: perfect alignment techniques might make the future ~20% better—the rest is misuse, inequality, chain-of-command, global stability, preference aggregation, meta-ethics, each needing its own intellectual community. Aligned shouldn’t be a synonym for good, Hay on alignment ≠ flourishing, and Hendrycks et al.’s Unsolved Problems in ML Safety are the same genre.
The important problems
| Problem | Why necessary | Earliest useful window | If late | Main dependencies |
|---|---|---|---|---|
| Control & verification | Everything else is moot | Before capability gates (now) | Irreversible loss of control / untrustworthy deploy | Not UBI; needs enforcement capacity |
| Bio & malicious use | Catastrophic misuse | Narrow legislative / industry window | Porous synthesis + models + cloud labs | Overlaps control; international arbitrage |
| Domestic governance | Firms won’t volunteer the alignment tax | State → federal; electoral cost | Path locked by industry | Unlocks distribution & HCI defaults |
| US–China / international | Unilateral pause and unilateral heavy tax fail | Summit and crisis windows | Race lock-in | Hard to tax alone either |
| Distribution | Alignment ≠ flourishing | Tax fix now; UBI late | Aligned but unequal (friction main dish in my forecast) | Needs governance; not ASI |
| Who writes the “constitution” | Aligned to whose values | Process design before lock-in | WEIRD / lab capture | Crosses governance & meaning |
| Human capability & meaning | Money ≠ living well | Product defaults now | Skill atrophy, anomie | Markets push automation |
Each problem below uses the same fields:
- Criticality — what happens if we don’t
- Key question — what’s actually stuck
- Status — where evidence and institutions are
- Timeline — which clock is running
- Good / bad forks — success / half-success / failure
- Under real constraints — friction, clocks, ROI; given that, what is more worth doing (A > B > C)—not “we should” slogans
Problem 1: Control and verification — don’t lose the wheel before “alignment”
Criticality
If frontier systems can’t be supervised, constrained, and checked after deployment, distribution and meaning are theater. This isn’t all of alignment research; it’s the engineering floor across the lifecycle.
Industry often collapses three different answers into “we’re aligned”:
| Layer | Question | Typical tools |
|---|---|---|
| Scalable oversight | How do weak supervisors train strong models? | RLHF, Constitutional AI, AI critics |
| AI control | Assuming the model may misbehave, how do we bound harm? | Sandboxes, permissions, monitors, vetoes |
| Verification | How do you know the first two worked? | Benchmarks, red teams, probes, logs, weight audits |
In oversight / control / verification I argued each layer can fail alone. A model can look aligned in training, pass deploy evals, and still scheme at runtime; a monitor can catch obvious sabotage and miss sandbagging under evaluation awareness.
Amodei et al. 2016 Concrete Problems in AI Safety listed five accident problems (side effects, reward hacking, scalable oversight, safe exploration, distributional shift)—ancestor of the problem-list genre, narrower than today’s control agenda. Hendrycks et al. Unsolved Problems in ML Safety expand to robustness / monitoring / alignment / systemic safety. The International AI Safety Report stresses an evaluation gap: pre-deploy tests often fail to predict real-world behavior.
Key question
As capability rises, can control stay competitive? Christiano’s old point: if alignment/control techniques aren’t competitive, markets pick stronger, riskier systems—you need global coordination to pay the alignment tax, or you need to drive the tax down.
Status
- Lab RSPs / safety frameworks are proliferating; most remain voluntary.
- External evals and red teams exist; “passed the benchmark” ≠ “won’t deceive after deploy.”
- Open weights: no recall once out; fine-tuning can strip refusals (IASR’s marginal-risk framing).
- In my forecast, durable pause is low-probability; compute monitoring is relatively more doable—detectable, quantifiable (GovAI compute-governance line).
Timeline
When 2026–2028 capability gates cluster, control that still lives in slide decks loses to deploy pressure. If interpretability becomes a real “MRI for AI,” that window is the same order of magnitude—Amodei has tied interpretability urgency to export-control timing.
Forks
| Picture | |
|---|---|
| Better | Oversight + control + external verification keep pace; high-risk capabilities default to sandbox + human veto; open weights get marginal-risk review |
| Half | Pretty benchmarks and RSPs; monitor weaker than model; “aligned” as PR |
| Bad | Irreversible deploy, then deception; or control tax so high nobody adopts it |
Under real constraints: research timelines, shelf tooling, who should do what
First, how far alignment research can realistically get—otherwise everything after is slogans. Numbers below are from my own forecast spine (U1 alignment & symbiosis; control side: oversight / control / verification, Node 11 corporate governance): rough judgment probabilities, not precision science:
| Item | Rough window / P | Read |
|---|---|---|
| Public “remote-worker-class” capability gate (C8-scale) | ~2027 H2–2028 H2 | Humans can still oversee, already strained |
| Internal superhuman research-assistant / alignment-scare window (C9–C10-scale) | ~2028 Q1–Q3 | Branch point; if info breaks, political window is short |
| Scalable-oversight-class methods hold through C8 | ~0.55 | Plausible, not locked |
| Nested oversight to C9 (≥2 levels) | ~0.38 | Transfer failure is the default risk (lab pretty, world collapses) |
| Interpretability artifacts actually gate ≥1 frontier release/year by 2028 | ~0.35 | Don’t treat “MRI for AI” as the modal path |
| Mandatory CoT + scheming classifiers on internal AI R&D by 2028 | ~0.42 | Measurable; buys time |
| Public RSP hard-stop invocation | ~0.12 | Mechanism exists; almost never used |
| ”Alignment symbiosis” all the way into long-run flourishing | ~0.18 (composite) | Needs many conditions at once |
Honest line: in 2027–2029, alignment research buys partial control and time—not “solved before C9.” Governance and distribution clocks will not wait for intent alignment to finish.
Second, what the industry can put on the shelf before policy arrives. Shock with nothing on the table is the FTX pattern (huge shock, still zero US federal crypto hard law). Shock with ready text and tools is what gets signed fast.
What is actually usable today splits in two layers:
- Public-goods eval stack: UK AISI Inspect (plus sandboxing toolkit), METR Hawk on Inspect; Apollo deprecated its internal framework for Inspect. In a crisis week, regulators copy shared eval/logging interfaces—not another open letter.
- Commercial security layer: Gray Swan–class products (Shade red-team, Cygnal runtime, Arena crowdsourcing; Series A–scale raise in 2026; cited across frontier system cards)—labs will pay, because it cuts PR, liability, and enterprise-sales friction. Lakera / CalypsoAI / Protect AI folding into legacy cyber vendors means this business is merging into cybersecurity, not birthing a separate “alignment industry.”
Safety NGO spend is roughly $200–400M/year against capability at something like 600:1. METR is on the order of a few dozen people. Under scarcity, ranking beats vibes.
Given that, more worth doing (A > B > C):
- A — Keep control/verification competitive as capability climbs: monitors, sandboxes, vetoes, evals that transfer (not just benchmark theater). Ctrl-Z–style results (attack success collapsing at ~5% capability tax in reported settings) point the right direction; generalization is still open.
- B — Inspect-compatible evals, incident schemas, compute-monitoring prototypes: the main field for researchers and independents. When policy arrives, regulators adopt ready protocols.
- C — Sell safety labs already buy (adversarial eval, runtime blocking): real commercial path; be clear it buys deployment misuse/jailbreak, weakly buys scheming/loss-of-control. Don’t market a firewall as “alignment solved.”
- D — Interpretability as one tool in a portfolio, not the whole bet.
- Low ROI: waiting for federal hard law before building shelf; proprietary eval stacks that don’t speak Inspect; another costless pause letter (FLI 2023: 1,000+ signatures, zero policy effect).
See also: alignment paradigms map · malicious-actor cost · oversight / control / verification
Problem 2: Bio and malicious use — not one slogan about bioweapons
Criticality
Catastrophic misuse doesn’t require uncontrolled superintelligence. It needs capability uplift + porous institutions. Synbio, protein-design tools, cloud labs, and frontier-model bio uplift together push the skill floor down.
Cotra puts misuse first among problems perfect alignment doesn’t solve: an AI aligned to a dictator is still aligned.
Key question
Can we install an institutional package before uplift is everywhere—not just model refusals?
Refusals get jailbroken and fine-tuned away on open weights. The real bottlenecks are physical and institutional: who can order synthetic DNA, who enters cloud labs, who gets high-risk biological AI tools.
Status: three layers, not one
A. Synthesis screening (DNA/RNA orders)
| Mechanism | Where it is | Holes |
|---|---|---|
| IGSC Harmonized Screening Protocol v3.0 | Members do sequence + customer screening; length thresholds tightening | Non-members, cross-border arbitrage, oligo pools, benchtop synthesizers, fragmented orders, novel sequences with no DB hit |
| US EO 14110 + OSTP nucleic-acid framework | Attestation / screening on federally funded paths | Not full commercial coverage |
| screendna.org open letter | Altman, Amodei, Hassabis, Suleyman + vendors + security community call for mandatory screening, recordkeeping, customer legitimacy | Not yet uniform US hard law |
B. Managed access for models
NTI bio’s managed-access framework for biological AI tools (Jan 2026): tier by tool risk; verify legitimacy; prefer API oversight over open weights for high risk; funders subsidize compute for compliant hosts.
“Know your scientist” three-tier KYC (institutional vouching → output screening → behavioral monitoring) can start without new statutes—but coverage equals whoever volunteers.
C. Cloud labs
Automated wet labs lower the hands-on bar. Few public biosecurity incidents yet; automation is still brittle—that’s a window, not permanent safety. RAND-style asks: a cloud-lab security consortium (IGSC analogue), KYC, human-in-the-loop for high-risk protocols, shared ban lists.
Timeline
“Mandatory screening this Congress” is the advocacy clock; wait past uplift and arbitrage plus benchtop gear tear the holes wider. In the bio evidence and science-abundance path, IGSC / mandatory-screening success multiplies onto “science abundance” paths.
Forks
| Picture | |
|---|---|
| Better | Mandatory synthesis screening + records (near-all vendors) + managed access for high-risk bio tools + cloud-lab KYC + international recognition (closes arbitrage) |
| Half | IGSC volunteers only + lab RSP refusals + open bio foundation-model weights → porous |
| Bad | Material uplift already diffused; screening still voluntary; benchtop bypass |
Under real constraints
Model refusals get stripped—that’s engineering fact, not pessimism. Real leverage sits in physical and institutional bottlenecks, and the window is before uplift is everywhere.
More worth doing (A > B > C):
- A — Push synthesis screening from voluntary to mandatory (federal statute + state/procurement pressure + screendna coalition growth) and close cross-border arbitrage. One of the few issues with bipartisan grammar and a physical chokepoint.
- B — Managed access for high-risk bio tools + “know your scientist” KYC: NTI’s framework is written; without mandates coverage equals volunteers—still better than refusals alone.
- C — Cloud-lab security consortium (KYC, human-in-the-loop for high-risk protocols, shared ban lists): few public incidents yet—that’s the window.
- D — Model-side CBRN evals and red teams: necessary, but a third door, not the only door.
- Low ROI: slogan “stop bioweapons” without touching orders/cloud labs; or waiting for “perfect alignment” before talking misuse (an AI aligned to a bad actor is still aligned).
See: p(doom) bio node evidence
Problem 3: Domestic governance — lots of paper, little conversion
Criticality
Without teeth: alignment spend won’t become industry norm (whoever pays the tax loses); redistribution won’t move unilaterally first; “leave room for humans” in HCI loses to pure cost-cut automation. Safety, economy, lived experience all need someone higher up setting and enforcing rules.
This won’t appear from nowhere. CEOs don’t lobby Congress to regulate themselves for mass unemployment.
Key question
Not “do politicians understand x-risk?” but: when does the political cost of inaction exceed the cost of action? 70%+ support regulation in polls; almost no federal hard law—the bottleneck is conversion. In making governments care and what happens before policy changes: thalidomide, EPA, secondhand smoke, Basel, EU AI Act—bills sit in drawers, shocks change intensity; FTX is the counterexample—huge shock, still zero US federal crypto statute.
Status (US, 2026)
Policy is written less on Capitol Hill than in statehouses and primaries. See US AI policy political landscape:
- California SB 53, New York RAISE: frontier transparency, incident reporting, whistleblower protections—first hard state templates.
- Federal: voluntary reviews, oscillating EOs; a “ten-year ban on state AI laws” style provision was killed 99–1 in the Senate.
- Elections: industry PACs spend hard to block regulation-friendly candidates (e.g. NY-12 Bores, ~$7M scale).
- At least seven factions—not a simple D vs R map.
Anderljung et al. Frontier AI Regulation compress needs into three blocks: standard-setting, registration/visibility, compliance (self-reg → mandate → license). The US is thinnest on federal compliance.
Timeline
2026–2028 election and admin cycles are themselves variables. State templates written before a crisis are what get signed fast when the window opens—Chatham House’s “ready-made” governance.
Forks
| Picture | |
|---|---|
| Better | State templates spread → federal visibility + compliance; independent eval capacity; transparent lobbying |
| Half | State patchwork, federal idle; voluntary RSPs sold as “already regulated” |
| Bad | Federal preemption kills state law with no substitute; or shock arrives with nothing on the table (FTX pattern) |
Under real constraints: money, elections, protest, whistleblowers
Lay out mid-2026 politics before “how to make governments care”—or it’s slogans again.
Federal: substantive frontier hard law still stalls; the executive track is voluntary review + preemption of state law. A December 2025 White House EO stood up an AI Litigation Task Force aimed at “onerous” state rules; Cruz-style “ten-year ban on state AI laws” died 99–1 in the Senate, but the preemption war continues (GAAIA-class drafts still float time-limited development-side preemption).
States: 100+ AI-related laws this term; SB 53 / RAISE are frontier templates. Policy is fought in statehouses and primaries, not finished on Capitol Hill.
Money: The AI Lobby–style stats: when opposition lobbying on a bill exceeds ~$10M, original-form passage ~12%; under ~$1M, ~62%. Industry super PACs (Leading the Future and orbit) and regulation-side PACs are buying elections. NY-12–scale ~$7M against a regulation-friendly candidate is mechanism, not anecdote.
Public opinion: 70%+ want regulation; almost no federal hard law—the bottleneck is conversion, not “educate the public one more time.” See making governments care.
Honest 2026–2028 expectation (flagged as speculative):
The modal path looks like state patchwork + federal soft/voluntary rules + preemption litigation, not a national pause. The 2026 midterms decide who writes the next state templates—often more real than “a safety-friendly White House in 2028.” Don’t make the presidency plan A. Labor shock, chatbot harm, and data-center siting may become ballot issues faster than pure x-risk; x-risk has to ride those coalitions, not wait for mass street mobilization.
How effective are protests?
PauseAI / Pull the Plug actions—even London’s early-2026 “largest yet”—are still hundreds, not tens of thousands (MIT Technology Review). FLI pause letter: 1,000+ signatures, zero policy effect. The Encode path (SB 1047 vetoed → SB 53 enacted) shows specific bills + state legislative entrepreneurship beat vague pause slogans. Street action helps—Overton, media, social license for inside players—but at current scale it cannot be the primary theory of change. My forecast already puts mass x-risk protest probability low; writing as if it’s coming is dishonest.
Whistleblower protection—institution, not hero story
California SB 53 (effective 2026-01-01) added a Labor Code chapter: anti-retaliation; large developers must offer anonymous internal channels; disclosure to the AG / federal authorities. The triggering “catastrophic risk” definition is hard: a single incident foreseeably and materially contributing to death or serious injury of more than 50 people, or more than $1 billion in property damage/loss, tied to specified model capability classes (bio attack, deceptive loss of control, etc.). Coverage is covered employees in defined roles, not everyone at the lab. OpenAI and Anthropic have published updated policies; a federal Grassley AI Whistleblower Protection Act (May 2025) would broaden to contractors and lower thresholds—timing uncertain.
The honest read: whistleblowing buys visibility and a political window, not a training stop. In forecast Node 4 (whistleblower / scare), the modal after a scare is continue training with ~10–15% marginal slowdown; durable multilateral pause is a few-percentage-point tail. Without shelf bills and electoral cost, disclosures dissipate—Saunders-to-Congress ≠ compute caps. Federal preemption that guts state duties can weaken the protections themselves.
Given that, more worth doing (A > B > C):
- A — Block/weaken federal preemption; keep state teeth; spread SB 53 / RAISE templates. Highest documented negative + positive lever combo. Factions and races: US AI policy political landscape.
- B — Encode-style bill entrepreneurship + staff education: signable drafts before the shock (thalidomide / EPA / EU AI Act pattern). Policy tracking and lobby maps are coalition infrastructure, not the end product. Conversion mechanics: making governments care.
- C — Elections and coalitions: fund state champions; child safety / deepfakes / labor / energy siting / x-risk cooperate bill-by-bill—don’t force one slogan.
- D — Legal defense and anti-retaliation infrastructure (counsel, equity-loss buffers, usable anonymous channels): make disclosure survivable; push federal expansion, but don’t make it plan A.
- E — Street protest: Overton and morale tool, explicitly after A–C.
- Low ROI: another costless open letter; betting the 2028 White House; treating “70% support regulation” as already won.
Problem 4: US–China and international — unilateral almost never works
Criticality
Something like nine-tenths of frontier compute sits on the US–China side. The EU AI Act can be complete and still limited if those two don’t coordinate. Unilateral pause picks a winner; unilateral heavy tax moves capital and compute.
Key question
Can we get a verifiable floor, not just summit communiqués?
Status: narrow deals + mismatched rooms
| When | Event | Substance | Follow-through |
|---|---|---|---|
| 2024-05 | First formal bilateral AI talks, Geneva | US: technical frontier risk; CN often tied to export controls | Limited progress—same room, different conversations |
| 2024-11 | Biden–Xi (APEC Lima) | Affirm human control over nuclear-use decisions; caution on military AI | Extremely narrow; no training caps, no verification. Reuters |
| 2026-05 | Trump–Xi Beijing-related statements | Talk of best practices / dialogue so non-state actors don’t get frontier models; US frames talks as possible because it leads | Narrowest common ground (non-state); state-vs-state race, caps, mutual audits mostly off table; chip licensing runs in parallel; verification still missing |
Deeper asymmetry: China’s two meanings of AI safety—Western catastrophic/alignment “safety” vs Chinese public “安全” (content, stability, controllable/trustworthy). Negotiators can say the same English phrase and mean different things.
The UN Governing AI for Humanity seven recommendations are an inclusive architecture, not an emergency brake on the US–China race.
Timeline
“We talk because we lead” cuts both ways: the window may close as capabilities converge. Non-state norms are the narrowest shared interest—terrorists, criminals, ransomware—a floor, not a ceiling.
Forks
| Picture | |
|---|---|
| Better | Real non-state access norms + mutual recognition on CBRN/synthesis screening + compute monitoring or eval mutual recognition; critical-infrastructure reliability standards (“fire codes”) |
| Half | Summit language + chip-deal narrative, no verification |
| Bad | Geneva mismatch repeats; export-control whiplash; race with no floor |
Under real constraints: a negotiation menu, not “promote dialogue”
A verifiable training-cap treaty is a few-percentage-point tail in my Node 9 (multilateral); forum soft law is modal. Export controls today serve China containment, not a global pause. The chip metering / proof-of-training needed for a compute pause is exactly the least mature verification layer. Under that reality, “promote US–China cooperation” without a menu is empty.
More worth pushing, in order (A > B > C):
- A — Concrete non-state access norms for frontier models (terrorists, criminals, ransomware): narrowest shared interest that might move. Make KYC and reporting checkable—not communiqué adjectives.
- B — Mutual recognition on CBRN / synthesis screening: technically discussable, trade-linkable, closes arbitrage—same leg as Problem 2.
- C — Narrow, verifiable eval and incident information sharing: far more realistic than training caps; feeds domestic governance and external eval.
- D — Compute-monitoring and verification R&D: diplomatic capacity-building, not something ready for a treaty.
- E — Training FLOP ceilings / multilateral pause: last; needs D mature + political scare. Selling this as the near-term main line is self-soothing.
- Low ROI: pretending Chinese and English “AI safety” mean the same thing; substituting export controls for safety talks; counting summit language as progress.
See: export-control politics · AI governance landscape · China’s two safeties
Problem 5: Distribution — after OpenResearch, stop hand-waving UBI
Criticality
Hay / Cotra / Acemoglu on one line: task alignment is the default equilibrium ≠ flourishing. Models that faithfully enrich shareholders are a compatible “alignment success”—the main dish in the friction region of my forecast.
USV’s AGI economy model: utopia and dystopia are two equilibria of one system; the difference is competition and redistribution. Redistribution without competition → monopoly prices eat the transfer.
Key question
Who pays, at what scale, who moves first politically—and what cash RCTs actually showed.
Status: state OpenResearch clearly
One of the largest US unconditional-cash RCTs (~$60M): Illinois and Texas; 1,000 low-income adults (ages 21–40) got $1,000/mo × 36 months vs $50/mo controls (n=2,000). Core NBER papers 2024–2025. Headline results:
| Moved | Didn’t / precise nulls |
|---|---|
| Basics spending ↑; leisure ↑ | Durable physical/mental health (can reject tiny effects) |
| Labor participation −3.9 pp; −1–2 hrs/week; ~−$2k/yr earned income | Job quality, net worth |
| Short-lived stress / food-security relief | Most child schooling; political participation |
| Medical utilization slightly ↑ | Human-capital investment overall null (younger → education signal) |
Four caveats before anyone extrapolates to “post-AGI UBI”:
- Not universal—selected low-income young adults.
- Three-year gift, not permanent.
- External donor money—no fiscal incidence (who pays).
- Amount far below wage-replacement post-scarcity scenarios.
Authors’ frame: cash buys flexibility/agency, not a substitute for health care, childcare, or housing supply. Altman later cooled on classic UBI toward ownership/compute-share rhetoric—consistent with the read: money necessary, not sufficient.
Fiscal scale: Nayebi threshold and SWF analogues
Nayebi (2025) gives a citable rough benchmark: under a conservative “no new jobs” setting, funding a UBI-sized transfer 11% of GDP ($12k/adult/year style) from AI-capital rents alone needs AI productivity on automatable tasks ~5–7× today’s automation at ~15% public capture Θ; raising Θ to ~33% halves the bar to ~3×. Past ~50% Θ, returns diminish, especially if oversight costs raise operating cost c. Market structure matters: monopoly rents lower the threshold; fierce competition raises it—again the USV tension: competition helps consumers and can make rent-funded UBI harder.
Analogues:
| Model | What | Cash to people? | Scale intuition |
|---|---|---|---|
| Alaska PFD | Oil → fund → annual cheque | Yes: $1,702 (2024); $1,000 (2025, often legislature-set) | Far below poverty/wage replacement; political football |
| Norway GPFG | Oil → global assets → budget | Mostly no direct UBI | Huge AUM; “silent” social dividend |
| OpenAI Industrial Policy for the Intelligence Age (2026) | National public wealth fund partly seeded by AI firms | Proposal | Lab-authored menu; not enacted |
Realistic ladder (don’t start at the UBI slogan)
- Fix tax bias favoring automation—US labor vs capital tax wedge (Acemoglu’s oft-cited ~25% vs ~5%). Undo distortion before inventing exotic taxes.
- Broaden the base: corporate/capital gains, automated labor, compute/power/token options; competition policy against monopoly.
- SWF or public equity—Alaska cheque vs Norway budget, whichever can survive politics.
- Then scale cash (UBI) or services (UBS). OpenResearch: cash ≠ health or meaning.
Timeline
Tax correction and state experiments can start now. Federal-scale UBI almost certainly late—after labor shock is a ballot issue, or after crisis. “Any country enacts a robot tax by end-2027” style markets stay modest; OECD BEPS took fifteen years for half a deal. “Politics won’t cooperate” isn’t a terminal fact—it’s a problem—but the honest line is: draft the bill before the crisis, don’t invent it in the rubble.
Forks
| Picture | |
|---|---|
| Better | Tax base tracks capitalization + live competition + SWF/equity share; cash or UBS has a payer |
| Half | White papers + Alaska-sized cheques while gains stay concentrated |
| Bad | Alignment succeeds, GDP looks great, most people’s bargaining power collapses |
Under real constraints: is big reskilling spend worth it?
“Spend on reeducation and job redesign” sounds more responsible than UBI—the evidence is colder than the slogan.
OpenResearch: cash buys flexibility and leisure; human-capital investment is overall null. Reading the UBI RCT as “retraining will happen automatically” is a misread.
US public scale: WIOA Dislocated Worker–class budgets are ~$1.1B/year; relative to “AI shock needs tenths of a percent of GDP” talk, that’s an order of magnitude short. In my education/reskilling crosscut: failure-by-2032 probability is high; a federal >$10B package is unlikely; retrain-beats-automation-clock is only ~0.25–0.30.
Trade Adjustment Assistance evaluations (Mathematica / Schochet et al.): participants get far more credentials and training; employment and earnings are often worse in the first two years post-layoff (people are in training, not at work); by year four the gap narrows, employment converges, earnings often still lag. Card-style ALMP reviews: training often small or negative short-run, better medium-run—that’s trade-shock era evidence, not a guarantee at AGI speed.
OECD: employer-funded AI-related training correlates with better self-reported outcomes—that’s incumbent job redesign, not mass reemployment magic. Formal catalogs with AI content are a few tenths to a few percent of courses; high-AI-exposure vacancies are far more common.
Read: reskilling isn’t zero; it also isn’t the antidote to “aligned but unequal.” Clocks: METR task-horizon doubling is measured in months; public retraining in years.
Given that, more worth doing on distribution (A > B > C):
- A — Fix tax bias favoring automation + competition policy: undo distortion, stop monopoly prices eating transfers—clearer ROI than a giant new training ministry (Acemoglu / USV same line).
- B — Employer-tied upskilling / role redesign: change task boundaries inside the same firm—closer to OECD evidence than “layoff then public bootcamp.”
- C — Targeted transition + wage insurance (TAA lesson: credentials ≠ wage recovery): for clearly displaced cohorts, income floor beats course titles.
- D — Fiscally funded cash or services (UBI/UBS): necessary floor; OpenResearch says pair with health/childcare/housing supply—don’t expect cash to become human capital by itself.
- E — Mass public “AI reeducation” as the main path: last; partial mitigator, not the flourishing narrative core.
- Low ROI: UBI slogans with no payer; reskilling slogans with no tax base or competition; Alaska-sized cheques as pretend solutions.
See: governing the machine economy · AI’s real economic impact · capitalism and inequality
Problem 6: Who writes the model’s “constitution” — preference aggregation
Criticality
When labs say a model is “aligned,” they usually mean it passed harm benchmarks, follows a system prompt, or scores well in preference comparisons. That’s technical alignment. The normative question: when stakeholders conflict, what should the system optimize? Doctor, family, HIPAA, patient—no universal right answer. Every training pipeline silently picks a moral structure.
In Cotra’s leftover list, preference aggregation and philosophy/meta-ethics are explicitly “more law and policy than one more ML experiment.”
Key question
Social choice (aggregating rankings into a collective choice) hits Arrow/Sen-style limits—mapped in impossibility theorems. RLHF’s Bradley-Terry assumes a transitive latent utility; if preferences cycle, the scalar reward is fiction.
Deliberative democrats (Landemore and others): legitimacy needs argument, not just vote tallies. AI can facilitate, translate, fact-check, cluster—it doesn’t automatically get authority to set which rights are off the ballot.
Status
| Practice | Who sets values | Limit |
|---|---|---|
| Lab Constitutional AI | Staff-drafted principles (often UDHR-inspired + product experience) | Developer overweight |
| RLHF | Annotator pairwise prefs | Social-choice compression; WEIRD samples |
| Collective Constitutional AI (Anthropic + CIP) | ~1,000 Americans via Polis → constitution → trained model | US sample; wiki-survey ≠ full citizens’ assembly; process knobs still lab-chosen |
A one-off PR “we asked a thousand people” is not a revisable, rights-floored, polycentric process. The Positive Alignment paper also pushes many centers of oversight—not one moral chokepoint.
Timeline
Value lock-in often happens when habits and infrastructure harden. Process design has to beat irreversible “default constitutions”—human norms drift; static RLHF labels risk lock-in.
Forks
| Picture | |
|---|---|
| Better | Transparent principle text; clear rules on who participates, what’s not votable (rights floor), revision cadence; community customization + thin global floor |
| Half | One Polis pilot sold as “already democratic” |
| Bad | Hidden system prompts as de facto constitution; or a single global moral office |
Under real constraints
Labs will not wait for philosophers to finish before shipping. Default “constitutions” are hardening via staff principles + RLHF labels. Collective Constitutional AI shows you can ask a thousand Americans—and also shows how easily a one-off PR consult gets sold as “already democratic.”
More worth doing (A > B > C):
- A — Write the process before lock-in: who participates, what’s off the ballot, revision cadence, how community customization sits on a thin global floor—the text itself is shelf inventory for hearings and crises.
- B — Many centers of oversight, not one moral chokepoint (Positive Alignment direction): more realistic than a global moral office, cleaner than pure lab fiat.
- C — Put social-choice limits into product and policy briefs: Arrow/Sen aren’t academic garnish; scalar rewards are fiction when preferences cycle—regulators and buyers need a one-pager they can hear.
- Low ROI: another one-shot “we asked the public” as the endpoint; or betting meta-ethics gets solved by philosophy first.
See: alignment paradigms map · universal values · impossibility theorems
Problem 7: Human capability and meaning — after the cheque clears
Criticality
Jahoda’s five latent functions of work—time structure, social contact, collective purpose, status, regular activity—money doesn’t buy. A large share of FIRE people return to work after financial independence—not for cash, for structure. OpenResearch’s health nulls and “flexibility” framing point the same way: cash is a floor, not a ceiling.
Markets default every task toward “AI does, human reviews”—passive use hurts self-efficacy; “human first, AI improves” preserves capability. In human–AI interaction models I use four modes: Generate / Scaffold / Challenge / Step Back. Wharton-style chess studies suggest system-regulated assistance beats user self-selection—people are bad at self-regulating.
Verification cost is also one reason headline AI advantages compress to real ~1–5×: humans judge, AI does mechanical work, verification gets cheaper. Designs that keep human capability and designs that maximize economic return are often the same design—until the market only optimizes cost per task.
Key question
How does “leave room for humans” stay commercially alive against full automation? Tax preference, regulation, procurement standards—or enough Klarna-style customer disasters?
Status
Product defaults are still Generate. Education and medicine have real augmentation cases and “one-click homework” hollowing. Meaning infrastructure—community, volunteering, creator economies, civic deliberation—has no off-the-shelf blueprint; meaning-institutions evidence and the Illich / Vallor / Graeber / Keynes line all say this is hard.
Timeline
Product defaults can change now. Waiting until a generation’s skills collapse to rediscover Scaffold mode costs an order of magnitude more.
Forks
| Picture | |
|---|---|
| Better | Scaffold/Challenge defaults; human judgment retained in high-stakes domains; shorter hours with non-work meaning structures |
| Half | UBI cheque + empty scroll time |
| Bad | Skill atrophy + anomie + aligned babysitter AI |
Under real constraints
Markets optimize cost per task; “leave room for humans” will not win by default. When mass public reeducation (Problem 5) can’t beat the capability clock, product defaults are a lever you can move now that directly affects whether skills atrophy.
More worth doing (A > B > C):
- A — Change product defaults: from “AI does, human reviews” toward Scaffold / Challenge; use procurement and licensing in high-stakes domains (medicine, education, law) to block pure Generate. One of the few things you can do now without waiting for federal UBI.
- B — Tax preference / procurement / liability so human-in-the-loop outcompetes full automation: Klarna-style customer disasters are how markets learn—too slow and too painful; rules can price it earlier.
- C — Employer-side role redesign (overlaps Problem 5’s B): change tasks inside the same organization—better for capability and meaning structure than retrain-after-layoff.
- D — Meaning infrastructure (community, volunteering, creators, civic deliberation): hard, slow, no blueprint—but OpenResearch and Jahoda both say cash alone isn’t enough. Long-term investment, not next year’s KPI.
- Low ROI: waiting until a generation’s skills collapse to rediscover Scaffold; treating “reeducation budgets” as the answer to meaning; defaulting to full automation then remediating afterward.
See: what makes humans happy · desirable difficulty · human–AI interaction
How the seven stack
flowchart TB
subgraph survival [Survival floor]
P1[Control and verification]
P2[Bio and malicious use]
end
subgraph rules [Rules layer]
P3[Domestic governance]
P4[US–China / international]
P6[Who writes the constitution]
end
subgraph flourish [Flourishing layer]
P5[Distribution]
P7[Capability and meaning]
end
P1 --> P3
P2 --> P3
P3 --> P5
P4 --> P5
P3 --> P7
P6 --> P1
P5 --> P7
Arrows aren’t a timeline—they’re dependencies: what blocks what. Control and bio can move first; redistribution is hard to force alone; without money or structure, meaning hangs in the air.
For what you can do by skill set, see how to participate in AI safety.
Sources
- Cotra, A. Aligned shouldn’t be a synonym for good. Planned Obsolescence.
- Hay. Solving alignment isn’t enough for a flourishing future. LessWrong.
- Hendrycks et al. Unsolved Problems in ML Safety. arXiv:2109.13916.
- Amodei et al. (2016). Concrete Problems in AI Safety. arXiv:1606.06565.
- Anderljung et al. Frontier AI Regulation. arXiv:2307.03718.
- UK AISI. Inspect. METR. Hawk.
- IGSC. screendna.org open letter. NTI. Managed access to biological AI tools (Jan 2026).
- California SB 53 (TFAIA); Labor Code catastrophic-risk whistleblower chapter.
- The AI Lobby — When Lobbying Spikes, Bills Die.
- MIT Technology Review (Mar 2026). London anti-AI protest.
- Reuters (Nov 2024). Biden–Xi: humans not AI control nuclear weapons.
- UN. Governing AI for Humanity.
- USV (2026). Modeling the AGI economy.
- Nayebi (2025). arXiv:2505.18687 (UBI / AI-capital rent thresholds).
- OpenResearch / NBER unconditional cash papers (2024–2025); Mathematica / Schochet et al. TAA evaluations (DOL).
- OECD. AI and skills / workplace AI surveys; Training Supply for the Green and AI Transitions (2024).
- Anthropic + CIP. Collective Constitutional AI.
- Positive Alignment. arXiv:2605.10310.
- White House (Dec 2025). EO on national AI policy framework / state preemption push.
Further reading (this site)
- My AI futures forecast — regions and probabilities · U1 evidence · futures evidence index
- Oversight / control / verification
- Making governments care · Before policy changes
- US AI policy landscape
- China’s two safeties
- Governing the machine economy · Human–AI interaction
- Alignment paradigms · Impossibility theorems