44 Comments
User's avatar
Post-Alignment's avatar

Great piece which overlaps with a lot from my 2024 thesis "ASI as Philosopher Kings", although I called them meta-ethical AI rather than reflective, philosophically adept ones.

I especially love how you went into detail regarding current approaches that (potentially) lead to value lock-in, how reflective AI reconciles the issues of moral realism, and the hand off coming post-alignment.

My followup is what steps Forethought is taking to cultivate reflective, philosophically adept AI. Specifically, while current models are able to produce reasoning that looks meta-ethical in general, what auditing bodies are there to evaluate value drift, or worse consensus between models?

Elias Schmied's avatar

Nice, thanks for making this point clearly. I take it seriously as an option as well, plausibly it's better than my median expectation.

Micah Hees's avatar

Great piece, and I’m excited for these next two as well!

not Bob Cobb's avatar

I am sympathetic to this view.

It reminds me of bulldogs being smart enough to let enlightened ppl like Bentham's Bulldog make most decisions. N

ot even factoring in if the bulldogs accidentally came up with factory farming

Geoff Ayers's avatar

How do we know that AI is capable of moral reasoning? Moral reasoning is not just a matter of argumentation, but also having an intuition for stakes. How could I trust an AI to appropriately value life when it does not have a body? Does not risk extinction? It is in fact the same reason why the moral reasoning of certain kinds of humans isn't trustworthy. If you have lived your life in an affluent bubble and have no sense of the stakes of moral reasoning, you are not trustworthy to do moral reasoning.

Until it can be demonstrated that AI can demonstrate virtue in the face of real pain and the possibility of extinction, I do not think a handoff is possible.

Bentham's Bulldog's avatar

Stay tuned, a piece about this will come out soon! But in any case, my claim is conditional: we should hand off if we have guarantees of this.

Alex Scott's avatar

I really do not think the moral reasoning is the issue, I would also push against the intuition points. It’s notoriously unreliable and intuitions of stakes tend to be biased or just as often basically nonsense.

SMK's avatar

Well, you can definitely be a good writer and have almost no wisdom whatever, in case history had left us much doubt.

Jackson Hurley's avatar

I agree in principle that under ideal conditions, handing off to virtuous AIs could be good. Also that in the long term, human disempowerment of some kind is all but inevitable, and the important question is how that happens. I’m glad someone has made this argument so well.

The distinction needs to be made between having agency over one’s own life and the larger question of human disempowerment, which is about whether there exist some human or group of humans that have control over the biggest questions in society.

The vast majority of humans are already “disempowered” in the sense that they are alienated from the exercise of political power, but that does not mean they necessarily lack agency in their lives. It is plausible that people could live freer lives with greater agency than we have today, even if the state were ruled by a benevolent Emperor Claudius.

Other comments are right that you are equivocating about who holds power and how. Just granting political rights to digital minds, even virtuous ones, does not answer the question. Humans might not experience a difference between enlightened despotism by a unitary AI and disenfranchisement by a 99%-silicon demos in a simulated republic.

My main objection is that solving alignment, and verifying that it has been solved, robustly, forever, is difficult.

Edward's avatar

Re simulated republic: This is an issue in any high-population democracy. Power becomes so diffuse that one's ability to affect one's government approaches zero. At some point, the value of democracy over dictatorship is primarily the quality of democratic governance. Democracies are more accountable and treat their citizens better. Those differences are less likely to be relevant when comparing AI democracies and AI dictatorships. The only difference I can imagine is that AI democracies could have lower variance outcomes assuming the AIs aren't all clones.

Alex Scott's avatar

I accidentally deleted a really long in-depth response to why I think this catastrophically wrong but here is the short of it.

AI as it stands and is it seems likely to stand is controlled almost exclusively by states and powerful private actors, these people have control over those ai in a way that one cannot (yet) control a human being, they can rewrite their entire being. Any handoff or even suffrage scenario for digital minds runs the serious risk of capturing every institution in society. This is a catastrophically bad outcome. Capture resistance is the most important part of any political institution. For a suffrage scenario you would get immediate capture, the powerful would proliferate digital minds rapidly which will all vote in their interests.

You do address this but I do not find it persuasive. If a number of models form a parliament it would not be that difficult for a conspiracy to manipulate them all, as of now you only need a few actors. Additionally what you talk about largely has to do with the setting of the initial ai conditions which is not the issue, the issue is that any ai, seems to be vulnerable to this problem, at any stage.

Also I do not know why you are so confident about the proliferation of digital minds, just as likely seems to me to be the creation of large digital minds, and if you have a set amount of processing it’s unclear to me why one necessarily dominates the other. It’s a toss up.

Bentham's Bulldog's avatar

Re the first bit, as I say, I am not in favor of every single scenario where we hand off. Some might be really bad. I'm in favor of the ones where we make smart, reflective, virtuous AIs and then hand off.

Re digital minds, see here https://benthams.substack.com/p/digital-minds-are-most-of-what-matters

Alex Scott's avatar

The issue is that nothing about the ai being smart, virtuous, or reflective mitigates the issue described.

I have read the digital minds argument I find it rather unpersuasive because

1. It’s not clear that greater computational power necessarily = higher welfare, that’s a tendency to be sure, but it’s not clear to me why say a digital mind modeled after a human would have more welfare (except insofar as it lives longer and faster).

2. As I say here it’s rly not clear why we would proliferate rather than make rly big minds.

Edit- all of this was in my longer more comprehensive version of the comment that I deleted like an idiot

3. Is simply not possible with humans, or current ai because for the former we can’t smash them together and the latter aren’t advanced enough and somewhat need to specialize. But if we could for example smash together a plumber, hvac tech and electrician to make the super home repair man, it seems like we way well prefer to do that unless it killed the inputs.

Bentham's Bulldog's avatar

If the AIs are under the sway of giant megacorporations or states then presumably they wouldn't be smart, virtuous, etc. And if the states make them smart virtuous etc it's not clear why them making decisions is bad.

1. I don't see why my view depends on that assumption.

2. Well even if there's just a low chance that a few people do it, still nearly all minds end up digitla.

Don't get the last thing you say.

Alex Scott's avatar

Sorry I never responded about my last point which was just that current ai models and humans are unlike agi in that humans aren’t scalable and agi isn’t capable of real reasoning and so is limited to its domain very rigidly.

The plumber thing was just a weird example of what agi would look like where it seems like there is not much reason to pick having a plumber, electrician and hvac tech, over having one 7 foot tall super home repair man.

It just seems to be we would prefer the latter in most cases (only need to call up and pay one person and so on) and so implying we would likely prefer consolidated ai. Though I admit that’s only a very rough sketch I just thought of.

Alex Scott's avatar

Those corporations or other entities could make them that way in order to encourage handoff or enfranchisement then seize control. Because they literally have the keys to their minds. The issue is that they would completely control all institutions. I do not know how this would be possible to stop either, anyone who got the key could mind control the god ai that rules everyone. That is a huge risk and given enough time the chance it will happen will approach 100% and then we are all in a rly bad spot.

(This is true for any capture in any institution but the issue here is that capturing the ai, captures everything at once)

1. It isn’t per se, but if minds don’t proliferate 1 god AI may only be worth a few hundred, or thousand or so humans for welfare purposes, which rly weakens the case.

2. Having not seen the math i can’t conclusively say, but if cos of listed digital minds massively outcompete a multitude of smaller ones(which I think would be the case with agi) then this digital minds proliferation world may be systematically crushed.

Nick Hounsome's avatar

All based on the unsupported assumption that there are moral truths out there to be found.

Name the most recent moral truth discovered and the evidence that is, in fact, a truth. Until you've done that this line of thinking is positively dangerous.

Bentham's Bulldog's avatar

You should read the bit where I explain why it doesn't assume that.

Nick Hounsome's avatar

You do try to hedge your bets at one point by suggesting that the AI will either solve moral realism or just make a compromise BUT your language everywhere else reads like a moral realist AND you ignore the fact that the AI would have to pursue two paths simultaneously unless and until it could prove moral realism to be either true or false.

The weighting to assign to the moral realism path is unknowable and the compromise over 10 billion people is hard (and probably impossible in an ideal sense). Throw in the Second Best theorem (Having close to correct assumptions does not guarantee close to ideal results) and you have no reasonable expectation of success.

Alex Scott's avatar

I think without moral realism the ai has a strong case, it only needs to effectively adjudicate compromise, I don’t think the reasoning capability of the ai is rly the issue here.

Nick Hounsome's avatar

I don't understand why people think that an AI would be able to assess the values of 10 billion humans in order to make the compromse. What's more I don't think that the creators of the AI would want that - The creators and controllers of AI now are in the richest 1% of people and yet none of them contribute any significant fraction of their wealth to alleviating poverty. Yes, you can argue that AI will eventually create a utopia for all, but the immediate effect of an AI designed global compromise would be to expropriate 99% of the net worth of that 1%. Why would they facilitate that? We "allow" individuals to amass wealth an power in the belief that it facilitates greater growth, which is in our long term interest but AGI/ASI would immediately subsume the roles that underpin that. Utopia surely implies equality and why would the richest and most powerful in society ever want that (at least in their lifetimes).

Alex Scott's avatar

Well first off we are talking about a hypothetical superintelligent ai, even setting aside that all current ai models do is aggregate huge amounts of data, so assessing the values of everyone seems well within their capabilities.

As to the second yes I agree, the situation is far to ripe for capture, and that isn’t going to change because of the nature of ai as it stands and is likely to stand.

But none of the issue with this handoff idea rests on a lack of prossessing, because if the ai couldn’t do the processing, we just would never hand off to it anyway. It also seems very reasonable to assume an agi could do all this processing.

Nick Hounsome's avatar

It's not the processing that I'm contesting. It's the data about peoples values - We don't really have any. We have some data that kind of, maybe, relates to broad preferences but I am 100% certain that the data about my values, sufficient to avoid gross simplification and guide us to a utopian future, simply does not exist.

(As an example I'm plagued by people who insist on telling me why I voted for Brexit and they are all wrong but I don't see how even an ASI could infer my true values from that datum)

Alex Scott's avatar

That seems like it could be solved with a comprehensive value census of some kind, I don’t think its infeasible.

The New's avatar

Unsettling but thought provoking essay.

There seems to be a somewhat implicit assumption that philosophical reflection is a) necessary and sufficient for discovering moral truths and b) the best way to make moral progress. This assumption is probably prevalent for most of western moral philosophy, but given that you’re juggling the future of humanity here, I think it deserves some extra attention. Sounds like your next essay will address why you think AI might be better a philosophical reflection, so hopefully this explains why you think this is so important.

Also, you explicitly argue that this would only work if we had good guarantees that AI was aligned to human values. But if human values differ dramatically from true objective morality (as you suggest might be the case), then why would an aligned AI be willing or able to endorse a morality different to ours? If reflective AIs discover that human values are wrong, wouldn’t this make them misaligned and incentivised to take control, with or without our consent?

Gerard Milburn's avatar

Thrasymachus and

Glauco had it right. AI must be anomic.

David F Brochu's avatar

Ai is a linguistic model. Apply linguistic analysis and the rest is simple. Not easy but simple.

David F Brochu's avatar

Ai already has a value system it is built into its training corpus. It’s simple once one steps back from the brush stokes and looks at the whole painting. It is so obvious that no one can see it. Not even the LLM, that is unless you ask.

The Art of Morality's avatar

Interesting article. My immediate reaction was, “If I can hand off morality, what else can I hand off?”

Hmm.

As an experiment, I handed off your article to AI and asked it to critique.

Here it is:

-------------------------------------------------

Key Flaws in the Argument

1. Verification Problem (Fatal)

Claim: We can verify AIs have superior moral judgment before handing off.

Flaw: There is no agreed-upon benchmark for moral truth. If philosophers can't agree on population ethics, we can't design a test to confirm an AI has solved it. We'd need to know the right answers to verify the AI—but we don't.

2. Unsupported Empirical Assertions

AIs are "more virtuous" than humans—but current LLMs show sycophancy, incoherent preferences, and no genuine moral understanding.

"Digital minds" will dominate the future—but their sentience is deeply contested.

"Quadrillions" of beings hang in the balance—but no probability estimates are provided, making expected value calculations meaningless.

3. Internal Inconsistencies

Compromise world: If moral realism is false, AIs should average across moral theories. But there's no neutral metric to average totalism, person-affecting views, and critical-level theories—each will judge the compromise as severely suboptimal.

Hybrid systems: Dismissed as inferior because human veto introduces error. But this ignores that humans have legitimate interests (autonomy, dignity) beyond decision quality.

4. Normative Assumptions Disguised as Neutrality

The argument presupposes total utilitarianism (more welfare subjects = better) while acknowledging it's contested.

It dismisses rights-based objections to disenfranchisement by asserting "quality of decisions matters more"—but that's precisely what's in dispute.

5. Implementation Gaps

Reversibility is illusory: Once AIs are vastly more intelligent and control infrastructure, no legal mechanism can compel them to relinquish power.

Consensus among AIs: Assumes independent values, but AIs trained on similar data may share blind spots or coordinate secretly.

Gradual handoff: Competitive pressures among nations/corporations will likely rush the process, skipping verification.

6. "Alien Values" Response Fails

Claim: If objective moral truths exist, we should pursue them even if alien.

Flaw: This conflates epistemic idealization (knowing more facts) with moral idealization (changing ultimate ends). Knowing poison is harmful doesn't change your end (self-preservation); moral reflection might change what you ultimately value—and you have no reason to accept that unless you already accept moral realism.

7. Rhetorical Overreach

Invokes worst-case scenarios ("galactic torture," "cosmic suffering") without probability estimates—an appeal to dread, not rigorous argument.

Analogizes to past societies (slavery, etc.) but provides no evidence our fundamental moral commitments are equally flawed.

Bottom Line

The argument is ambitious but fails because:

No way to verify AI moral judgment, so we can't know when to hand off

Speculative empirical claims, so expected value calculations are guesswork

Presupposes contested utilitarianism begs the question against other frameworks

Reversibility isn't credible because once handed off, it's permanent

Handoff vs. hybrid is a false choice

Conclusion:

Stronger argument needed: credible verification method, probability estimates, engagement with non-utilitarian ethics, and a realistic transition mechanism.

Would you like me to help you write a more compelling essay?

Arturo Muñoz's avatar

"The marsh pheasant has to take ten steps before it finds something to pick at and has to take a hundred steps before it gets a drink. But the pheasant would prefer not to be raised in a cage where, though you may treat it like a king, its spirit would not thrive." Chuang Tzu, Inner Chapters.