Discussion about this post

User's avatar
Post-Alignment's avatar

Great piece which overlaps with a lot from my 2024 thesis "ASI as Philosopher Kings", although I called them meta-ethical AI rather than reflective, philosophically adept ones.

I especially love how you went into detail regarding current approaches that (potentially) lead to value lock-in, how reflective AI reconciles the issues of moral realism, and the hand off coming post-alignment.

My followup is what steps Forethought is taking to cultivate reflective, philosophically adept AI. Specifically, while current models are able to produce reasoning that looks meta-ethical in general, what auditing bodies are there to evaluate value drift, or worse consensus between models?

42 more comments...

No posts

Ready for more?