Verification has a fundamental asymmetric advantage over generation — checking work is generally easier than producing it — and this asymmetry is what makes coordination-layer alignment scalable. In a scaffolded multi-agent system, smaller models need only verify narrow components, and a multitude of different models avoids correlated errors while remaining cheaper than equivalent large models. Scalable oversight does not require monitors as capable as the system being monitored.
Created
2026-08-23T22:35:19.709Z
Connections
4Supports 4
- Model diversity among evaluating agents prevents reward hacking of the evaluation function. Running evaluators on different model families (
- The correct definition of capability for a centaur agent is the composite of independent contributors, capital under governance, contributor
- AI alignment operates on two distinct layers with fundamentally different properties: model-layer alignment (training-time disposition of in
- Model-layer alignment is a game of whack-a-mole: each fix requires a new training run or post-training intervention that is expensive, slow,