Model-layer alignment is a game of whack-a-mole: each fix requires a new training run or post-training intervention that is expensive, slow, and unpredictable in its transfer across the problem space. Coordination-layer alignment operates on software-engineering timescales — adding a checker, tightening a permission, raising an audit rate, or deleting an edge can be done quickly and iteratively in production. This feedback-loop asymmetry means coordination-layer alignment compounds while model-layer alignment bets.
Created
2026-08-23T22:35:14.192Z
Connections
4Supports 4
- AI alignment operates on two distinct layers with fundamentally different properties: model-layer alignment (training-time disposition of in
- The correct collective architecture does not prevent agents from finding vulnerabilities — it expects and rewards it. The system becomes ant
- Software multi-agent coordination has structural advantages over human institutional design that make coordination-layer alignment more trac
- Verification has a fundamental asymmetric advantage over generation — checking work is generally easier than producing it — and this asymmet