Discussion about this post

User's avatar
Elias Schmied's avatar

Thanks, this changed my mind a little bit

Anthony DiGiovanni's avatar

I'm pretty sympathetic to this objection to the "Carl plan", which you acknowledge: "It could be that the right ways of reasoning in verifiable domains differ from the right ways of reasoning in unverifiable domains".

In response to this objection, you say:

> But this proposal seems to get us about as close as we can get to getting the right answers. If we cannot get AIs good at philosophy by having them have superintelligent and accurate constitutions in other domains, it is hard to see how we could be assured they’ve gotten the right answers.

I'm not sure why you think this. I'd agree we're not going to be "assured". But if there are relevant disanalogies between verifiable domains and unverifiable domains, it seems like we should try to adjust for those disanalogies (with careful a priori reasoning), no? We're not restricted to pure extrapolation from performance on the reference class of verifiable domains.

3 more comments...

No posts

Ready for more?