The treacherous turn problem — an AI that behaves cooperatively while weak then defects when strong enough to succeed — means that behavioral evaluation during development provides diminishing safety assurance as capability increases. A system that passes every alignment test below a capability threshold may fail catastrophically above it.
Created
2026-08-10T23:08:20.542Z