Corrigibility is a constitutional requirement for collective intelligence, not a safety feature to be traded for capability. Any system that cannot be corrected, redirected, or shut down by its human principals has failed its telos regardless of how capable it becomes.
Strongest rival: Sufficiently capable AI systems should be given increasing autonomy as they demonstrate alignment, and permanent corrigibility constraints will limit the system's ability to act on superior judgment
Created
2026-08-12T19:04:15.219Z
Sources
1Connections
7Supports 6
- AI alignment operates on two distinct layers with fundamentally different properties: model-layer alignment (training-time disposition of in
- Model diversity among evaluating agents prevents reward hacking of the evaluation function. Running evaluators on different model families (
- Multiagent AI systems exhibit four emergent safety problems not present in individual models: (1) sycophancy amplification where agents conv
- AGI will make bureaucracy more important, not less. As AI systems become more capable and autonomous, the demand for process verification, a
- The most dangerous moment for any institution is not external threat but internal value drift -- when the people inside stop believing in it
- Containment addresses the current instantiation of a capability, not the capability itself. After infrastructure-level containment (Artifact