Alignment is a living property
Every proposed fix for AI safety has the same shape: a property tested once, in a system that keeps changing.1 Train the model, evaluate it, ship it. Then the world moves, the model is fine-tuned, a new tool is attached, and yesterday’s safety case quietly expires.
Model-layer fixes are whack-a-mole. Each failure prompts a new training run that is expensive, slow and specific to the failure that caused it.2 Worse, once a benchmark becomes the target it stops measuring what it was built to measure.
Alignment is not an end state you reach. It is something the whole human–AI system does, continuously, or stops doing.3 That moves the question from “is this model aligned?” to “is this system still correctable?”
That is a design problem, and designs can be inspected. The rest of this part argues that a collective of differentiated agents, with attributed memory and humans in the loop, is the most inspectable design we have.