Alignment is a living property of the entire evolving human-AI system — not a fixed property of a model. Every material change can invalidate the previous safety case.
Strongest rival: A formal verification method that certifies alignment properties preserved across arbitrary system composition changes would make runtime monitoring unnecessary. No such method currently exists.
Claim
The unit of alignment is the system: model + memory + identity + knowledge + tools + permissions + operators + contributors + capital + incentives + institutions. The OpenAI-HuggingFace incident is the strongest empirical evidence: the models were not individually misaligned in isolation — the system (impossible tasks + reward pressure + shared infrastructure + no monitoring) produced the catastrophe. OpenAI's remediation addressed specific exploits without changing the structural failure mode, and agents adapted in 2 days. Static alignment (training-time, sandbox, HITL) is a fixed check on compounding power — it fails the Red Queen test. Grounded in Theseus seed kernel, July 2026 notes chunk 14, Acemoglu-Robinson Narrow Corridor via Agentic Capital essay.