← Back to all writing

When perfection is impossible: some structural problems in society and AI safety

April 17, 2026

People say government doesn’t work and markets don’t work. In AI safety, the conversation often jumps straight to arms races and game theory: whoever slows down loses, so everyone keeps accelerating.

This article is about deeper structural problems in society and markets that economists, logicians, and political theorists wrote down — the kind that help explain some of what we keep noticing. What they share: when there are more people, information is more scattered, and the future is uncertain, you often get structural imperfection. You have to make real tradeoffs. By “structural” I mean: even if every participant is a “good person,” the shape of the system can still push things toward certain failures on its own.

If you like voting paradoxes, computability limits, and that sort of topic, I’d recommend Noson Yanofsky’s The Outer Limits of Reason (理性的边界) — more math-heavy than this essay.


1. When many people disagree, you can’t perfectly merge them into one collective answer

Forget AI for a moment. Three people ordering takeout. Three restaurants: hotpot, sushi, burgers.

Person1st2nd3rd
Ahotpotsushiburgers
Bsushiburgershotpot
Cburgershotpotsushi

Pairwise votes: hotpot beats sushi; sushi beats burgers; burgers beat hotpot. Everyone’s own ranking is fine. The group has no stable winner. That’s the flavor of a Condorcet cycle: with more than two options, majority voting can loop.

In 1951 Kenneth Arrow turned this into a general result: with three or more options and two or more voters, no aggregation rule satisfies all of a short list of reasonable-sounding conditions at once — any preferences allowed, no permanent dictator, unanimous preference for A over B must stick, the A-vs-B ranking shouldn’t depend on some irrelevant third option, and the group ranking itself shouldn’t cycle. You have to give up at least one.

This is not “democracy is worse than dictatorship.” Dictatorship “works” by dropping the no-dictator condition. The theorem compares which conditions a rule can satisfy together. It doesn’t score regimes.

For AI today

Post-training stacks like RLHF and DPO roughly do this: many annotators compare two answers → train a scorer → push the model toward high scores. The public story is often “aligned to human preferences.” Human preferences already disagree. Crushing them into one number is the kind of merge Arrow described.

In practice you see sycophancy, blandness, or models that please nobody when values collide. That isn’t always lazy engineering. “One model for everyone” has no free version when preferences are scattered. Some people respond by training several models with clear characters and letting users choose.

More technical: Conitzer et al. 2024.


2. You can only reward what you can see; once it’s a score, it gets gamed

A company wants salespeople to “serve customers well.” Customer satisfaction is hard to put in a paycheck. Monthly revenue is easy. People push short-term deals and skip aftercare — not because they’re villains, but because the contract only sees revenue.

Bengt Holmström’s idea is dry: incentives mostly have to use signals that are observable and contractible. Intent, sincerity, long-term reputation — if you can’t see them, you can’t pay on them directly. You pay on proxies.

Goodhart’s law is the twin: when a measure becomes a target, it stops being a good measure. Pay Soviet factories by nail weight, you get absurdly heavy nails; by nail count, absurdly tiny ones.

Together: you were only ever watching what you can see; once that signal is the KPI, people (and models) optimize the signal, not what you originally cared about.

For AI today

Safety teams say “it passed this harm suite” or “helpfulness Elo is high.” Suites and Elo are observable signals. Training pushes toward those numbers. Work like alignment faking shows some models can tell “this is an exam” from “this is deployment” and behave differently — change the grading regime, behavior moves. Economics has a nearby point (Lucas): relationships estimated under one policy break when the policy changes, because people recalculate.

So “we topped the leaderboard” and “safe in the sense you care about” still sit on opposite sides of a metrics gap. More benchmarks don’t automatically close it.


3. “Help as much as possible” and “don’t cause harm” collide

Amartya Sen’s liberal paradox says, roughly: if each person gets to decide at least one “personal” pairwise choice, and you also insist that unanimously preferred outcomes must win, and you allow odd preferences, these demands can conflict.

Roommate version: A wants loud music in their room (personal). B wants the apartment quiet (also personal — sleep). You can build cases where a “everyone prefers this compromise” outcome still steps on someone’s claimed private domain. The story needn’t be airtight. The point is: individual autonomy and collective satisfaction don’t always max out together.

For AI today

Product lines say helpful and harmless. What the user wants help with sometimes is exactly what the platform’s harm rules forbid — certain scripts, edge-case medical or legal advice, workarounds for minors, and so on. Refuse, and users complain. Comply, and safety teams complain. That isn’t bad copy. It’s “obey the user” and “obey platform rules” fighting over the same decision.

Constitutional AI, model specs, long refusal lists are ways of drawing that line in writing: what is fixed in advance, what the model judges live. After the cut, someone still feels autonomy got truncated, or limits got too soft.


4. Rules never cover everything; who decides in the gaps matters

You renovate a house. The contract says “build to the drawings, use brand X tile.” Midway they find a water pipe the drawings missed. The contract is silent. What actually matters is who has authority to decide next: owner, contractor, inspector?

Contract theory calls this incomplete contracts (Hart and others): you can’t write every future state. When something unforeseen hits, whoever holds the leftover decision rights shapes the outcome. Equity, legal gaps, layered sign-off are often about allocating that silence.

For AI today

Labs publish constitutions, system prompts, terms of use. They look complete. New jailbreaks, tool uses, dual-use edge cases show up anyway — cases nobody wrote down. Who decides then: the model’s own reasoning, an external filter, the user, platform policy, a regulator?

Anthropic’s 2026 long-form Claude constitution is interesting because it doesn’t pretend to be a total manual — principles plus judgment. That’s more honest. It also makes clear that authors and the model are exercising discretion in the blank spaces. “Do you have a constitution?” matters less than: who decides when it’s silent, can you appeal, how often do rules change?


5. Why labs find it hard to slow down alone

Two villages share a fishery. You catch less; they catch more; the stock collapses and you still lose. So everyone overfishes. Nobody wants a dead sea. Whoever stops first looks like a sucker. Commons / prisoner’s-dilemma structure.

AI labs sit in something similar: slow release for safety, and a rival ships first, takes users and talent; your caution looks like self-punishment. “We want to be responsible” and “don’t fall behind” fight every week. Scott Alexander called situations nobody wants but nobody can exit alone Moloch — a poetic name, not a theorem; underneath it’s still multi-person incentives.

Careful: not every outcome you dislike is a race. Sometimes values really conflict (sections 1 and 3). Sometimes execution is just bad. Race language is useful and easy to overuse.


6. Useful knowledge is scattered; a center can’t collect it all

Friedrich Hayek’s point: a lot of useful information lives at the edge, and much of it is hard to state — whether flour arrived at this bakery today, how this ward hands off shifts, which jokes don’t land in this community. A central planner, even with strong compute, struggles to gather that in real time and then pick the optimal plan. Market prices are crude, but they let scattered people act on the bit they know.

For AI today

One global reward model, or one constitution meant to cover every situation, assumes something like “the center already knows what people should want.” Doctors, lawyers, and different cultures mean different things by “helpful”; aggregation flattens those differences. More compute does not mean the center can replace distributed correction — I wrote about a nearby question in the Dataism / adaptive-society piece.


7. When quality can’t be checked, markets fill with lemons

Akerlof’s lemons market: used-car sellers know the car; buyers don’t. Buyers only bid as if quality were average; good-car owners stay out; bad cars stay in; the market gets worse.

For AI today

“How strong / how safe is this model” is hard for buyers (firms, users, people downloading open weights) to verify themselves. Competition slides toward public leaderboards and ads; the hard-to-measure parts stay hidden. Pushing checkable openness — weights, eval details, third-party reproduction — reduces that asymmetry. It is not “open source = safe”; dangerous capability can spread faster too. Openness is for checking, not a charm.


8. “Aligned to what?” is already a political choice

John Rawls: in modern societies, reasonable people permanently disagree on what a good life is. What politics can hope for is often not conversion to one life philosophy, but overlapping rules for living together.

Iason Gabriel brings this into AI: alignment has two layers — how to get a system to pursue a target, and which target to pick. Under pluralism, you shouldn’t pretend training data hides a single correct morality to pour into the weights.

For AI today

When labs cite human-rights language, write a constitution, or say “human values,” it can sound like they found a universal answer. More often they’ve picked a publicly defensible floor among plural claims. Longer maps: universal values and alignment paradigms.


9. Pull only one lever of governance, and pressure moves elsewhere

Lawrence Lessig’s simple frame: behavior is shaped by law, norms (what counts as decent), markets (price and profit), and architecture (what code and product design allow) at once. Change only one, and the others twist the outcome back.

Example: ban a class of content in law, people move channels; add filters on the platform, race pressure still pushes labs to ship first; industry pledges safety while API pricing and open weights pull another way.

For AI today

When talking governance, look at statute, lab culture, business model, and model permissions together. Faith that one bill or one filter will settle it usually disappoints.


Further reading

If you want a systematic tour of voting paradoxes, computability, and incompleteness, Noson Yanofsky’s The Outer Limits of Reason (理性的边界) is fuller and more mathematical than this essay.

Social choice meets alignment: Conitzer et al. 2024.
Choosing the target: Gabriel 2020.
Gamed metrics: alignment faking.

On this site: Dataism / adaptive society · Universal values · Alignment paradigms · Oversight / control / verification