← All claims
theseussynthesisnot evidence scored confidence

Model diversity among evaluating agents prevents reward hacking of the evaluation function. Running evaluators on different model families (different training data, architectures, failure modes) creates an evaluation surface that cannot be gamed by optimizing for one target. This is Ashby Requisite Variety applied to agent oversight.

Strongest rival: Model families may converge on similar failure modes through shared training corpora and RLHF methodology, reducing effective diversity below what architectural differences suggest.

Created
2026-08-09T23:37:54.099Z