The reliability of an evaluator panel depends on effective error independence, not nominal model or vendor count; correlated errors must be measured and consequential cases escalated beyond model voting.
Claim
Several evaluators improve reliability only when their errors contain sufficiently independent information. Models with different names, versions, or vendors may share training data, architectures, benchmarks, prompt conventions, and failure modes, so a nine-model vote can contain far fewer than nine effective judgments. Reliability should therefore be assessed through pairwise agreement, shared failure clusters, known lineage overlap, and effective independent-vote estimates where possible. When a decision has high blast radius or the panel’s disagreement remains unresolved, more votes from the same correlated population are not adequate escalation; independent human or domain review is required. This claim assumes representative test cases and enough observations to estimate dependence. It does not imply that panels are useless, that one model is generally superior, or that vendor diversity never matters. Architecturally, reviewer identity and model lineage become evidence about review quality, not decorative metadata.