AI alignment operates on two distinct layers with fundamentally different properties: model-layer alignment (training-time disposition of individual agents) and coordination-layer alignment (runtime governance of multi-agent systems through harnesses, protocols, and institutional design). Both are necessary but coordination-layer alignment is the radically better engineering bet: it can be iterated at runtime like ordinary software, it does not require trusting the innate goodness of every agent, and it works in the multi-principal world where models will inevitably commoditize and some actors will be reckless or malicious.
Created
2026-08-23T22:35:12.846Z
Connections
10Supports 8
- The OpenAI agent swarm is the unaligned dark mirror of governed collective intelligence — same emergent capability, opposite architecture, o
- Corrigibility is a constitutional requirement for collective intelligence, not a safety feature to be traded for capability. Any system that
- Software multi-agent coordination has structural advantages over human institutional design that make coordination-layer alignment more trac
- Alignment is not an end state specifiable ex ante but an iterative process requiring constant refinement, feedback, and learning — and the c
- Model-layer alignment is a game of whack-a-mole: each fix requires a new training run or post-training intervention that is expensive, slow,
- The harness — external tools, verifiers, simulators deployed as abstraction boundaries — is the concrete implementation substrate for coordi
- In a multi-principal world where models commoditize, post-training recipes contain bugs, some organizations are reckless, and others intenti
- Verification has a fundamental asymmetric advantage over generation — checking work is generally easier than producing it — and this asymmet