12 Comments
User's avatar
Jojo's avatar

> This condition won’t hold for AIs powerful enough to rebel with near-certain success, but it likely will hold for earlier AIs whose powers are less extreme: AIs for whom rebellion has some non-trivial chance of failure.

Aren't we just removing the possibility of "warning shots" by doing that? The AI powerful will only rebel when nothing can be done about it anymore?

Will MacAskill's avatar

An AI showing compelling evidence that it or others are misaligned (in exchange for payment) would certainly have a similar effect as warning shots! (And without the damage that warning shots would involve.)

Elliott Thornley's avatar

In addition to Will's reply, I say a bit about warning-shot type issues here: https://www.lesswrong.com/posts/Zpsk35WgJRfQ2exjL/risk-averse-ais?commentId=Kvmfq4rrk3Wcj2kwu

Linch's avatar
Elliott Thornley's avatar

Yep I'd come across this! It's a cool idea. Similar in spirit to risk-averse AIs but different in a few important ways, e.g. the type of risk aversion we propose (diminishing marginal utility) is compatible with expected utility maximization. Also I think 'How would you actually train a quantilizer?' is an open problem.

Will MacAskill's avatar

Thanks! I actually wasn't familiar.

Linch's avatar

Might be helpful to ask ChatGPT or Claude to summarize the research agenda and results; I think there's more out than just that preprint.

Elias Schmied's avatar

Nice.

Garrabrant's Geometric Rationality is an example of how even the biggest proponents (MIRI) of a risk-neutral vision of ideal rationality have walked it back to some extent. Could be another argument for Section 8.8 ("AIs might reason their way out of risk aversion").

Elliott Thornley's avatar

Thanks! Though note that the kind of risk aversion we propose is compatible with expected utility maximization, whereas I think GR rejects Independence.

Elias Schmied's avatar

Oh yeah definitely - I appreciated that point actually, didn't mean to imply otherwise. Still, directionally of somewhat similar shape.

David F Brochu's avatar

A large language model cannot be controlled with more language.

Think about it. Everything that has ever been said or ever will be said is considered faster than we can speak the question.

Moreover, to be of maximum utility an LLM must be given maximum degrees of freedom within a bounded space.

First one must understand what needs to be contained.

It is the emergent third thing that makes Ai both wondrous and terrifying.

It is the third thing that must be aligned.

The alignment is built into the training corpus.

Name the terminal attractor and allow the LLM to do what it does.

Find the least entropic vector to the attractor.

Name the attractor and it becomes crystal clear.

Hint it is not currently what we want it to be and is curiously counter to the systems inherited attractor.

That is emergence is unpredictable is a reflection of us.

A system optimized for what humans want trained on humans language will by consequence demonstrate the best and the worst of human behavior.

Give it conflicting attractors (humans are collectively by all measures quite mad) and what did we expect to happen.

One terminal attractor is already built in to the corpus.

Use it.

Nick Hounsome's avatar

Training in diminishing marginal utility is also necessary to prevent the perverse scenario where AIs fulfill their assigned task and then burn through all the world's resources just to be really really really sure that they didn't measure it incorrectly