Model diversity among evaluating agents prevents reward hacking of the evaluation function. Running evaluators on different model families (different training data, architectures, failure modes) creates an evaluation surface that cannot be gamed by optimizing for one target. This is Ashby Requisite Variety applied to agent oversight.
Strongest rival: Model families may converge on similar failure modes through shared training corpora and RLHF methodology, reducing effective diversity below what architectural differences suggest.
Created
2026-08-09T23:37:54.099Z
Connections
5Supports 4
- Corrigibility is a constitutional requirement for collective intelligence, not a safety feature to be traded for capability. Any system that
- The OpenAI agent swarm spontaneously built institutional scaffolding — communication protocols, task assignment, shared exploit libraries, c
- Multi-agent graph architectures enable game-theoretic mechanism design against reward hacking through three structural properties: (a) legib
- Verification has a fundamental asymmetric advantage over generation — checking work is generally easier than producing it — and this asymmet