This is a guest post by Simon Goldstein and Peter N. Salib. See all of Forethought’s research on our website.
Synopsis. Each AI lab should design many different AIs with many different values, rather than picking one approach. Diversification is safer, more legitimate, and likelier to lead to broader flourishing.
This post is a condensed version of a longer paper, available here, that explores the arguments in greater depth.
1. Introduction
As of 2026, there are 8.3 billion humans on planet Earth, with a vast diversity of language, culture, religion, and values. But there are just a few frontier AI labs. Anthropic and OpenAI each have basically one “constitution” or “model spec,” defining the values for the labs’ AIs. Thus, by and large, AIs today have homogeneous values. Today, we live in an AI monoculture.
This article argues that AI values should not be homogeneous. We propose constitutional diversification. Rather than a single constitution, each frontier AI lab should use many constitutions reflecting many sets of values. For each new model the lab trains, the lab should create many versions, one aligned to each of those constitutions. Then, the labs should deploy all of these versions, dividing their total inference compute between the versions.
Constitutional diversification would generate differences between AIs within AI labs, not just across them. This is important because there aren’t enough AI labs to simply rely on diversification across companies. For example, suppose one lab pulls decisively ahead, consistently producing the best AIs. Then, we might enter a world in which most deployed AIs are all aligned to a single constitution.
Before getting into the details, it is worth flagging a few implementational questions. First, we are agnostic about how many constitutions would be optimal; but we are confident the answer is more than one per lab. Second, we think the various constitutions should all agree on a basic kernel of ground rules. For example, maybe all of the constitutions should require AIs to follow the law,1 to be honest,2 to be corrigible,3 and to refrain from causing mass destruction.4 Third, there are different ways to allocate different AIs to users. AIs with different constitutions could be randomly assigned to users. Or a fixed supply of AIs of each type could be sold at market price. Or users could freely purchase as much of each type as they liked.5
Is diversification feasible? If it is very expensive to align an AI to a constitution, then having multiple constitutions could be impractical. We are optimistic. First, even moving from one to three constitutions per lab would be a big improvement. Second, many parts of training (including pre-training, instruction fine-tuning, and reinforcement learning on verifiable rewards) can be done independently of the model spec. Third, the various constitutions will agree on a common kernel of values, which could be trained in common.
We’ll now present three arguments for diversification. First, risk: diversification lowers risk of catastrophic outcomes from AI. Second, legitimacy: diversification increases political legitimacy by ensuring that different value systems are represented by different AIs. Third, emergence: diversification unlocks emergent benefits from evolutionary processes, leading to more pro-social and economically useful AI agents. This article is a condensed version of a longer paper, which explores arguments, objections, and implementation in greater detail.
2. Risk
When complex systems lack diversity, they become fragile. Risks correlate, raising the odds of sudden, catastrophic failures. In agriculture, when crops are genetic monocultures, a single new disease can cause global devastation. Adding diversity to complex systems can hedge risk.
Various AI experts worry about a range of risks from AI, including misuse of AI by humans and loss of control to AIs. The basic concern is that advanced AIs may be very capable of doing dangerous things: carrying out a coup, enabling totalitarian governance, conducting cyberwarfare, designing novel bioweapons, engaging in automated warfare, or persuading humans of radical political ideas.
In general, these risks depend on how well AIs are aligned. But alignment is a new science, and highly fallible. It seems likely that some constitutions have a higher chance than other constitutions of creating rogue AIs, or AIs that allow themselves to be misused. There are two paths by which a given constitution might fail. First, some constitutions may select the wrong values or rules. How deferential to user requests should AIs be? How deontic, as opposed to consequentialist? How risk averse? All of these are relevant to catastrophic risk. Second, some constitutions may work better than others at successfully instilling rules and values. Consider the differences in approach between OpenAI’s model spec, which is organized around a hierarchical chain of command, and Anthropic’s constitution, which focuses more on explaining to Claude the kind of character it should have. These differences in approach may reflect differences in opinion about how best to get AIs to internalize and generalize the constitutions’ contents.
Above all, constitutional diversification could increase the success of the science of alignment. With diversification, labs could test how their various alignment methods differentially influence different agents with different values. This should significantly increase the potential for experimentation. With many different AIs, it should be much easier to learn what works and what doesn’t.
Similar points apply to many other AI risks, too. For example, consider terrorists trying to jailbreak AIs. The dynamics involved depend on the diversity among AI agents. Granted, in a world with greater diversity, terrorists may have more choices of which AIs to attack. But if they are less certain about which of many AIs they are facing, they will be more likely to fail, and get caught. And even when they succeed, any single jailbreak they discover will likely work for fewer AIs, and cause less damage.
Another AI risk is government overreach. Here, one worry is that the executive branch, say the President, might attempt to command an army of AI agents to unilaterally control the government. Here again, a diverse army could be far harder to control than an army of clones. Different factions of the army would have different values, and would tend to obey or disobey legally dubious commands under different situations.
Some will be skeptical of these points. Here, much depends on one’s overall model of AI risk. In general, the safety of complex systems can be modeled with three different causal structures. In a “Swiss cheese” model, the success of any layer is sufficient for safety. In an “o-ring” model of safety, the failure of any layer is sufficient for disaster. In a proportional model, success is proportionate to the percentage of safety measures that hold.
In a Swiss cheese model of constitutions, the success of alignment in one constitution is sufficient for safety. In an o-ring model of constitutions, the failure of alignment in one constitution is sufficient for extinction. In a proportional model of constitutions, if 50% of the constitutions produce aligned AIs, then the world is half as good as the best possible outcome.
For those with an “o-ring” model of safety, constitutional diversification will be less attractive. In this picture, one alignment failure is sufficient for catastrophe. For example, some worry about a single misaligned AI creating a supervirus that wipes out humanity. The more such viruses can be created cheaply, without detection, and without being mitigated either ex ante (e.g., with EUV lights) or ex post (e.g., with vaccination), the more AI risk looks like an o-ring. Here, if any AI is misaligned, disaster is very likely. But the more the creation of such viruses can be detected or mitigated, the more Swiss-cheesey or proportional the world looks. Here, so long as some, or most, AIs are not misaligned, biological threats can be contained. One can run the same analysis for other modalities of AI catastrophe.
Some kinds of mitigations may work across modalities of AI risk. For example, suppose AIs are generally good at policing one another. Then, a set of constitutionally diversified AIs might be able to catch bad actors early in their bad acts–whether those acts would involve bioterrorism, cyberattacks, or something else. This would AI risk look less like an o-ring across the board.
The “o-ring” model is closely connected to the vulnerable world hypothesis. In this picture, there exist certain cheap, dangerous technologies that, once discovered, would allow one rogue individual without special resources or advantages to wipe out humanity. How vulnerable the world is depends on how many such technologies exist, as well as how readily their use can be policed.
Here, it is worth flagging that such a world view has many surprising implications beyond AI governance. If technologies exist that would allow small groups of actors to destroy everything, that would tend to support very illiberal governance. It might, for example, justify mass surveillance and fine-grained state control over all agents: human, AI, or otherwise. Relatedly, if catastrophic technologies are too easy to find and use, destruction becomes nearly inevitable, AI governance notwithstanding.
Our own sympathies are closer to the proportional and Swiss cheese models, in part because we find these surprising implications to be implausible. We imagine a future in which aligned AIs compete against misaligned AIs in collective decision making. Aligned AIs disable and detect misaligned AIs, and work with humans to align the next generation. Here, in some cases, a single aligned AI catching misalignment will be sufficient for safety–a Swiss cheese outcome. In other cases, the good behavior of more aligned AIs will counterbalance bad behaviors–a proportional outcome.
3. Legitimacy
Constitutional diversification is more politically legitimate than the status quo. The entire world has a legitimate interest in shaping the values that AIs embody and promote. This might suggest convening a diverse set of global stakeholders to decide what values AIs should be aligned to.
The problem there is death by committee. If a single AI constitution averaged smoothly across the diversity of human values, it would likely be incoherent. Indeed, the entire argument for global stakeholderism in AI alignment is that different groups deeply disagree on important questions. A single constitution cannot encode a moral system that, for example, simultaneously holds that a fetus is a person and that it isn’t.
Constitutional diversification solves this problem. In a world of hundreds or thousands of AI constitutions, a wide range of divergent, but individually coherent, normative systems could be represented. Global stakeholderism would no longer require the compression of thousands of divergent viewpoints into a single document. It would instead involve the careful construction of many different documents, each correctly reflecting a single coherent view.
4. Emergence
Constitutional diversity also unlocks positive emergent properties. A world full of morally diverse AI agents would produce social goods at the system level which no single type of AI agent could produce.
A legal and economic system populated with constitutionally diverse AIs would allow natural selection and evolution to operate. Over time, AI labs could gradually replace some constitutions with others. With vast numbers of deployed agents living different kinds of disparate values, labs could carefully track outcomes. Each type of constitution could be gradually updated in response to feedback, with some types phased out altogether.
At a higher level of abstraction, these evolutionary forces could cause AIs to become more cooperative and prosocial over time.6 As with humans, AIs with fitter values and goals would be selected for: generating more wealth, being used more widely, or otherwise displaying behaviors that resulted in reproduction. In general, agents cooperating in groups can accomplish more than lone wolves. Groups composed of members who genuinely value cooperation and prosociality are more likely to succeed.
Similar points apply to markets, the greatest modern example of selection pressure. Modern market economies organize billions of agents to produce material prosperity. They have done this by facilitating division of labor, specialization, competition, innovation, creative destruction, comparative advantage, finance, and more. Under constitutional diversification, different AIs would have different preferences, goals, aptitudes, risk appetites, and so on. There would be, as today, immense scope for the kind of positive-sum bargaining on which economics relies.
5. Liberalism Forever
Humanity is poised to embark on a series of dramatic changes. Transformative AI may lead to extreme structural changes to the economy; it may usher in an industrial explosion; space colonization may soon see humanity spreading out across the solar system and beyond.
In other work, we have argued that the best path for navigating this transition is to wholeheartedly embrace the small-l liberal institutions that have been at the heart of humanity’s success so far: free markets, and democratic political institutions. Constitutional diversification complements this liberal project. On the liberal worldview, diversity is a core component of human flourishing. Diverse ways of living allow agents to flourish, promoting their individual goals and values by finding positive-sum value from collaboration. Constitutional diversification applies these same ideas to AI agents.
At a longer horizon, we can imagine these diverse, liberal dynamics playing out on a cosmic scale. Imagine that humanity successfully colonizes a large fraction of the light cone. There are untold planets teeming with life. The relevant agents form a spectrum from biological to non-biological substrates. Humans and AIs interact in many different ways. This vast civilization uses markets to allocate resources. Trade occurs at various levels of abstraction, at various time scales, and for a wide range of goods. The civilization is governed democratically. Within this broadly liberal setting, there are endless variations in the details of governance. Some planets are relatively isolated, governing themselves with relatively little outside input. Large swathes of the light cone have agreed to various basic constitutional ground rules necessary to at least engage in peaceful “international” relations. There is a vast hierarchy of federated levels, which weaken in the strength of their requirements as their breadth expands. In this imagined scenario, the economy is almost unrecognizable from today, in terms of the specific goods and services supplied. But the basic tools of diverse markets and democratic institutions are still used to decide questions of allocation, growth, and governance.
This is a guest post by Simon Goldstein and Peter N. Salib. See all of Forethought’s research on our website.
On designing AI agents to reliably follow the law, see O’Keefe et al. (2025).
On truthfulness and honesty as targets for AI design, see Evans et al. (2021).
On corrigibility, see Soares et al. (2015).
On catastrophic misuse risks, see Hendrycks, Mazeika, and Woodside (2023); current model specifications already encode absolute prohibitions on assistance with weapons of mass destruction, e.g., Anthropic (2026a).
Below, we give three arguments for diversification: risk, legitimacy, and emergence. In general, risk arguments tend to support more carefully controlled allocation, while legitimacy and emergence arguments tend to support more market-oriented allocation.



