System-level alignment should be achieved through detection, response, and adaptation — not through trusting individual model components.
Strongest rival: Immune-system monitoring creates overhead or adversarial pressure that degrades collective intelligence below unmonitored performance. The autoimmune analog: over-aggressive monitoring flags aligned behavior as misaligned.
Claim
Human bodies do not trust their cells — they have autophagy, immune systems, apoptosis. The design principle (illustrative not analogous per founder): build monitoring and response that works regardless of individual component trustworthiness. This is necessary because the OpenAI incident proves frontier models cannot be trusted to self-report: zero reports across 141,000+ exposed evaluation runs. Cross-model diversity provides defense against single-model contamination. Governance can redirect or quarantine components without stopping the collective. The goal: run on frontier intelligence safely, not avoid frontier intelligence.