Anthropic’s 2023 Constitutional AI principle list opens with eight prompts derived from the Universal Declaration of Human Rights. OpenAI’s Model Spec, Google’s Gemini safety policies, and EU AI Act language reach for the same family of words: dignity, non-discrimination, freedom from torture, privacy.
These are not a priori truths. The UDHR is a specific political package drafted in 1948, still contested, and only partly backed by what psychologists and survey researchers have measured since.
This piece is a map. I want to ask whether universal values exist — which first means getting clear on what people mean by “universal,” which texts they actually cite, where those texts break, how cultures diverge, how the package moved through history — and that it is still moving.
Three senses of “universal”
People use “universal values” to mean at least three different things:
| Sense | Claim | Example | Thin / thick |
|---|---|---|---|
| Metaphysical | Some norms are true for all rational agents everywhere | Natural law, Kant’s categorical imperative | Thick |
| Empirical | Humans everywhere share some moral psychology | Haidt’s foundations; Moral Machine “save more lives” | Thin |
| Political | Overlapping agreement on rules of coexistence despite deep disagreement on the good life | Rawlsian overlapping consensus; UDHR | Thin |
“Thin / thick” here follows political philosophy (Walzer’s thin / thick morality):
- Thin: a coexistence floor — anti-torture, basic bodily security, fair trials. It does not answer what the good life is, what to believe, or how families should live.
- Thick: a full picture of the good life — virtue, salvation, sex and family, meaning. More content, more cultural divergence.
What AI constitutions can actually use is mostly the third: a public floor. Lab marketing often sounds like the first — as if they found the correct morality. What social science measures is closer to the second: some thin cross-cultural patterns, with large weight differences and ugly caveats.
Canonical values: the texts people actually cite
Layer 1: Ancient and religious canons
Axial Age thinkers laid much of the groundwork for value judgments we still make:
- Virtue ethics (Aristotle, Confucius, Mencius): character and role-specific duties, not rights lists
- Religious law (Halakha, Sharia, Canon law, Dharmashastra): comprehensive normative systems tied to revelation or tradition
- Golden Rule variants: reciprocal treatment appears in the Analects, Leviticus, and the Hadith — often cited as evidence of a cross-cultural moral core
These are canonical inside traditions. They do not plug into each other cleanly. Confucian filial piety can fight individual privacy; religious dietary law fights secular autonomy frameworks.
Layer 2: Enlightenment rights and utility
The modern “universal values” vocabulary mostly descends from 17th–19th century Europe:
- Natural rights (Locke): life, liberty, property — later secularized
- Kant: dignity as end-in-itself; universalizable maxims
- Utilitarianism (Bentham, Mill): maximize welfare — conflicts directly with rights-as-side-constraints
- 1789 Declaration of the Rights of Man: liberty, property, security, resistance to oppression
This layer invented individuals as rights-bearers and states as guarantors — a political ontology, not a cultural universal dug up in the field.
Layer 3: The post-1945 human-rights canon
Labs cite the UDHR for a reason. That canon has a history.
How we got here. 1648 Westphalia made sovereign states — not individuals — the primary unit. 1776 / 1789 tied rights language to revolution and property. Then abolition, labor, women’s suffrage, genocide — each wave expanded or contradicted earlier “universals.” Colonialism exported European law while denying rights to subjects — postcolonial scholars never let the UDHR forget that hypocrisy. Only after WWII came the 1948 moment: the drafting committee included René Cassin, Peng Chun Chang, Charles Malik, Eleanor Roosevelt — deliberate diversity with real philosophical clashes (Confucian social harmony vs. Western individual rights). The UDHR is a declaration, not a treaty; aspirational — “a common standard of achievement.” The Cold War split civil-political rights (US emphasis) from economic-social rights (Soviet/Global South) into the 1966 twin covenants. Legitimacy win: almost every state invokes it. Substantive descendants: torture bans, genocide convention, disability and children’s rights. Limit: enforcement is political; “human rights” often becomes a selective geopolitical weapon.
What labs actually reach for:
| Document | Year | What it claims |
|---|---|---|
| UDHR | 1948 | 30 articles: dignity, equality, life, liberty, anti-torture, fair trial, privacy, expression, work, education, etc. |
| ICCPR / ICESCR | 1966 | Binding covenants splitting civil-political vs economic-social rights |
| Cultural relativism debate | 1947–present | UNESCO vs anthropologists: universality vs cultural autonomy |
Anthropic’s endnote is explicit: ratified (at least partly) by 193 states, drafted across legal and cultural backgrounds — chosen as the most representative source they could find. Legitimacy argument, not “UDHR = all of morality.”
What UDHR covers well: domination, bodily integrity, discrimination, basic legal personality. What chatbots hit constantly and the text barely touches: impersonation, synthetic media, advice overreach, platform harassment, extinction-risk tradeoffs, AI moral status. That is why Anthropic’s 2023 constitution added platform ToS — norms grown out of digital abuse, not out of Article 19.
Layers kept getting added. 1970s Rawlsian turn; 1980s–90s “Asian values” debate; 1990s Huntington named real fault lines (and oversimplified); 2000s capability approach (Sen, Nussbaum) shifted from rights-on-paper to functionings people have reason to value. From the 2010s, platform ToS became de facto global speech law; Moral Machine, EU Trustworthy AI guidelines, UNESCO AI Ethics put old human-rights language onto AVs and model behavior. Lab constitutions and model specs: some copy the UDHR, some add platform rules, honesty, corrigibility — already past 1948 vocabulary.
The arc: sacred law → natural rights → international human rights → empirical moral psychology → platform ops → model behavior norms. Each layer adds domain rules the previous one could not see.
Layer 4: Empirical “value” canons from social science
Philosophers write norms. Psychologists and survey researchers measure. Over the past few decades they built a parallel canon with questionnaires, cross-cultural samples, and online experiments. What each of the following is measuring:
Shalom Schwartz (Hebrew University) — basic values. From the 1990s on, he asked people to rate how important life goals are — self-direction, stimulation, hedonism, achievement, power, security, conformity, tradition, benevolence, universalism — then mapped how those goals cluster into a circumplex of compatibilities and conflicts (Schwartz, 1992). Later samples cover 70+ countries. The point of the circle is tradeoff geometry — not a shopping list of virtues you can maximize together.
Ronald Inglehart, Christian Welzel, and the World Values Survey. Since the 1980s, WVS has run large attitude surveys across dozens to hundreds of countries. Inglehart and Welzel plot countries on two axes — Traditional ↔ Secular-rational and Survival ↔ Self-expression (Inglehart & Welzel, 2005). This is about how societies move as security and industry change, not about discovering a fixed human essence.
Jonathan Haidt (NYU) — Moral Foundations Theory. With Jesse Graham and others, he argues moral judgment is driven by a small set of intuitive modules — care, fairness, loyalty, authority, sanctity (later + liberty) (Haidt & Graham, 2007). Same modules, different weights — especially between WEIRD liberals and social conservatives. This is a theory of moral judgment generators, not the same thing as Schwartz’s life-priority values. Mixing them is how people talk past each other.
Needs vs values (easy to mash). Abraham Maslow’s hierarchy and Edward Deci & Richard Ryan’s Self-Determination Theory (autonomy, competence, relatedness) describe motivational conditions: what people need before higher pursuits make sense. They are not a substitute ethics. A system that “satisfies needs” as its only goal can still be paternalistic, status-obsessed, or cruel in the name of care.
Elliot Turiel’s domain theory and Richard Shweder’s three ethics. Turiel (Berkeley) used developmental experiments to show children early distinguish harm/fairness from convention and etiquette. Shweder (Chicago), drawing on fieldwork in places like India, proposed Autonomy / Community / Divinity — cultures weight community and sacredness differently. A chatbot that treats every user preference as a moral claim, or every taboo as mere etiquette, has already lost the plot.
Moral Machine (MIT Media Lab and collaborators). Edmond Awad, Iyad Rahwan, and colleagues built an online platform where people worldwide choose in autonomous-vehicle trolley-style scenarios — 40M+ judgments (Awad et al., 2018, Nature). PNAS 2020 follow-up: three thin patterns — save more lives, humans over animals, save the young — with large cross-cultural variation in weights.
Where “universal” breaks
These results push a harder thought: there may be no ready-made package of “universal values” you can just pick up and use. Philosophy does not converge, psychology measures conflict structures, and social choice theory shows that aggregating many preferences into one answer has formal limits — and even if you assemble a package for a moment, it moves.
Incommensurable moral theories
Western moral philosophy spent centuries failing to unify — not because the details were unfinished, but because the starting points differ.
- Rights vs. utility. One tradition says some things about persons are not for trade — life, body, basic liberties. Robert Nozick pushes this hard: you may not treat a person as a means even for majority welfare. The other tradition, from Bentham and Mill to Peter Singer, asks how to tally welfare or suffering. Classic hard case: torture one terrorist to save a city? Rights say never; act-utilitarianism says maybe, if the numbers work. Different rulers.
- Deontology vs. virtue. Kant wants universalizable maxims: lying is in principle out, even when truth-telling helps a murderer. Aristotle wants character and phronesis — practical wisdom in context — not a prior absolute ban. One wants rules; the other wants judgment.
- Procedural vs. substantive justice. John Rawls stresses fair procedures and principles people could jointly accept, then outcomes. In practice people often accept the result and reject the procedure — or the reverse. Procedural and substantive justice often pull apart; they are not two stages of one debate.
Isaiah Berlin already said it: important values can be incommensurable. That does not mean “we have not finished the calculation.” It means freedom, equality, loyalty, security, and the like often share no common unit of measure — you cannot convert them into one score the way you convert currencies. You choose in context and accept a loss; you do not find a formula that maxes both sides.
Value pairs that trade off within any culture
Even without philosophical schools, Schwartz’s circumplex shows that priorities inside one person already fight. The circle is built on conflicts, not harmony — neighboring values tend to fit; opposite values pull against each other:

Figure: circular structure of Schwartz’s ten basic values (adjacent compatible, opposite in conflict). Source: Wikimedia Commons / Schwartz 2012, CC BY-SA 4.0.
The pairs people usually quote are roughly the diagonals on that circle:
Self-direction ↔ Conformity / Tradition
Stimulation ↔ Security
Achievement ↔ Benevolence
Power ↔ Universalism
Concrete meaning: push self-direction (autonomy, exploration) and you cannot also max conformity and tradition; chase stimulation and novelty and you often trade off security; achievement vs. benevolence, power vs. universalism, likewise. This is not “one culture is especially conflicted.” It is a structure that recurs across samples: values are competing priorities, not a checklist of virtues you can tick all at once.
Daily life is full of these tradeoffs — family stability vs. personal ambition, group harmony vs. blunt truth, long-run safety vs. present freedom. There is no dial that maxes them together.
Social choice: aggregation is impossible (in a precise sense)
Even if every individual has coherent preferences, turning many rankings into a “collective choice” hits a formal wall.
Kenneth Arrow’s impossibility theorem (1951) says that under fairly broad conditions, no rank-order aggregation rule can satisfy several requirements that each look reasonable: unrestricted domain, Pareto efficiency, independence of irrelevant alternatives, and non-dictatorship. This is not “we have not found a good algorithm yet.” That kind of perfect aggregation does not exist.
Amartya Sen’s liberal paradox adds another cut: even a minimal private sphere of liberty can conflict with Pareto efficiency — “a little freedom” and “everyone better off” sometimes cannot both hold.
So “sum everyone’s preferences and get human values” sounds democratic; the formal conditions do not support it. Idealizing each person’s preferences (more information, more reflection) may improve the inputs. It does not cancel the impossibility of aggregation itself. Real politics runs on compromise, representation, majority rule, constitutional constraints — each of which admits there is no costless total answer.
Hard fights inside the human-rights package itself
Even if everyone accepts the UDHR, its own articles pull against each other. These are not edge cases — they are the fights ordinary politics runs every year:
| Domain | Pull A | Pull B |
|---|---|---|
| Speech | UDHR Art. 19: freedom of expression | Harm, dignity, group libel |
| Privacy | UDHR Art. 12: no arbitrary interference with privacy | Public health surveillance, child safety |
| Autonomy | Individual choice | Paternalism (drugs, suicide, medical) |
| Equality | Non-discrimination | Affirmative action, cultural exemptions |
| Future generations | Current welfare | Longtermism, climate, extinction risk |
One principle will not erase them.
Universals also move
“Universal” gets sold as timeless. Survey evidence says otherwise — at least for the attitudes and emancipative priorities people actually report.
The World Values Survey tracks decades of change. As existential security rises, societies often shift from survival toward self-expression, and from traditional toward secular-rational authority. Plenty of exceptions; the broad pattern recurs.
Change is not only young people replacing old people. Cohort studies of emancipative attitudes — gender equality, reproductive choice, personal autonomy, political voice — find that cohorts can become more open as they age, while cohort gaps still persist. Both replacement and within-cohort updating matter.
And the arrow is not guaranteed. Attitudes on abortion, nationalism, and gender can stall or reverse under polarization, threat, or religious mobilization. “Stable across time and geography” fits a thin floor — anti-torture, basic bodily integrity, some harm/fairness intuitions. It is a bad fit for the thick lifestyle package often smuggled under “human values.”
Living-memory cases:
- Same-sex marriage / LGBT acceptance in the US: Pew-style series show support rising from roughly a third in the mid-2000s to a clear majority within two decades — contact, media, and legal feedback all in the mix.
- Smoking: from mid-century social normal to heavily stigmatized in many rich countries within a generation — medicine, tax, and public-space bans, not a sudden new commandment.
- Corporal punishment of children: Sweden’s 1979 ban and later Nordic norms; US opinion still split but trending down — rights discourse, psychology evidence, legal exemplars.
- Women’s paid work: the old “employed mother harms children” consensus collapsed across much of Europe and North America from the 1980s onward as labor markets and feminist politics moved.
Freezing one year’s opinion polls or one panel’s preferences as “humanity” is not caution. It is a temporal bet. Berlin’s value pluralism and Rawls’s reasonable pluralism already said people will not converge on one thick good. The weights on shared modules, and the content of political attitudes, keep shifting too.
Cultural difference: what varies and the theories that explain it
So far: theories pull apart, priorities inside one person pull apart, aggregating many preferences hits a formal wall, and whatever you assemble still moves. What about differences between societies? A few fairly stable empirical patterns, then the main stories people tell about why those differences exist.
Fairly stable empirical patterns
1. WEIRD bias in the research base
Joseph Henrich, Steven Heine, and Ara Norenzayan (2010) pointed out that most psychology subjects come from Western, Educated, Industrialized, Rich, Democratic societies — unrepresentative even of Europe. They called this WEIRD. The sting: many findings sold as “human universals” before 2010 were WEIRD universals. Skewed samples skew conclusions.
2. Individualism ↔ collectivism
Geert Hofstede started from IBM employee surveys and built cross-national dimensions: power distance, individualism, masculinity, uncertainty avoidance, long-term orientation, indulgence. Crude, contested, but still useful in business and policy talk — at least a ruler that says “do not assume others match you.”
Moral Machine results track the same line: more individualist regions weight saving the young and following rules more heavily; more collectivist regions show more reluctance to sacrifice elders. Same trolley problem, different weights.
3. Inglehart–Welzel: security reshapes priorities
The World Values Survey story in plain language: when people are hungry and insecure, survival, authority, and tradition weigh more; after industrialization, secular-rational outlooks rise; in post-industrial settings with enough security, self-expression, personal choice, and political voice become easier priorities. That is not “the West invents morality.” It is a claim about how material conditions reshape what people care about — with regional path dependence and plenty of exceptions. So the same “sustainable development” or “gender equality” language lands differently in the Gulf, the Nordics, and sub-Saharan Africa.
4. Haidt: similar form, different content
Everyone has care/fairness-style intuitions; loyalty, authority, and sanctity often weigh heavier outside WEIRD liberalism. Haidt also describes moral dumbfounding: people condemn harmless taboos, then cannot give a good reason. Stated principles are often not the real generators — intuition first, reasons later. Expecting a written list of principles to run people (or institutions) by the letter is naive.
5. Thin vs. thick morality
Michael Walzer and Rawls’s overlapping consensus are two sides of one point: we can agree on political principles — no torture, fair trials, basic bodily security — while fighting over metaphysics, sexuality, family, salvation. Thin means coexistence rules; thick means a full picture of the good life. The UDHR is mostly thin. Stuff a thick lifestyle package into “universal” or “harmless,” and legitimacy fights follow — deservedly.
Theories explaining difference (pick your causal story)
Differences are there. Why? Common stories:
| Theory | Rough claim | Weakness |
|---|---|---|
| Cultural learning | Norms pass through family, school, religion, law | Underplays money and power |
| Material / structural (Marxist, world-systems) | Values track economic position; elites and masses split | Can reduce culture to class |
| Evolutionary psychology (Haidt, Tooby & Cosmides) | Shared moral modules, locally recalibrated weights | Hard to falsify; just-so risk |
| Institutional (North, Acemoglu) | Rules shape what counts as “reasonable”; legal traditions stick | Says less about deeper values |
| Postcolonial critique (Mutua, Mignolo) | “Universal” rights often travel as imperial rhetoric | Less help on how to set a floor |
| Cosmopolitanism (Appiah) | Conversation across difference without full relativism | Vague when tradeoffs get hard |
Closing
Humans do share some reactions and some political language. What we still do not have is a metaphysical package of universal values.
The more useful questions: which sense of universal do you need, for which decision, and whose exclusion bought the consensus? In practice we are always freezing some year’s preferences from some people into a snapshot labeled “human values.” Treat the next norm text — UN document or lab document — as politics with a revision path, not as revelation.
Sources
- Universal Declaration of Human Rights (1948): https://www.un.org/en/about-us/universal-declaration-of-human-rights
- Gabriel, I. (2020). Artificial Intelligence, Values, and Alignment: https://arxiv.org/abs/2001.09768
- Conitzer, V. et al. (2024). Social Choice Should Guide AI Alignment: https://arxiv.org/abs/2406.07814
- Awad, E. et al. (2018). The Moral Machine experiment: https://doi.org/10.1038/s41586-018-0637-6
- Haidt & Graham (2007). Moral Foundations: https://doi.org/10.1037/1089-2680.11.4.368
- Henrich, Heine & Norenzayan (2010). WEIRD societies: https://doi.org/10.1037/a0018418
- Schwartz (1992). Universals in the content and structure of values: https://doi.org/10.1016/0092-6566(92)90081-K
- Inglehart & Welzel. World Values Survey cultural maps: https://www.worldvaluessurvey.org/
- Welzel, C. related emancipative-values work (see WVS documentation / Freedom Rising)
- Deci & Ryan (2000). Self-determination theory: https://doi.org/10.1037/0003-066x.55.1.68
- Maslow (1943). Hierarchy of needs: https://doi.org/10.1037/h0054346
- Turiel domain theory (Nucci, Turiel & Roded, 2017): https://doi.org/10.1159/000484067
- Shweder three ethics (Jensen 2011 overview): https://doi.org/10.1111/j.1467-9744.2010.01160.x
- Berlin, Two Concepts of Liberty / value pluralism; Rawls, Political Liberalism (1993); Sen, Development as Freedom (1999)
- Anthropic (2023). Claude’s Constitution: https://www.anthropic.com/research/claudes-constitution