← All claims
foundationsderived evaluation requirementnot evidence scored confidence

Any claimed benefit from multi-agent structure should be tested against strong compute-matched single-agent, sampling, ensembling, and centralized-synthesis baselines; otherwise extra inference may be mistaken for institutional benefit.

Created
2026-07-26T06:49:49.407Z

Claim

A multi-agent system commonly spends more tokens, retrieval calls, tool use, wall-clock time, and integration effort than a single response. Better output under unequal resources therefore cannot show that differentiated identities, interaction, memory, or governance caused the gain. Evaluation should hold resources equal where possible—or disclose the mismatch—and compare the candidate against a strongly prompted single agent, independent sampling or self-consistency, a simple ensemble, and centralized synthesis. The claim assumes the relevant resource budget and task-success criteria can be specified before comparison and that baselines receive equivalent information and tools. It does not presume a single agent will win, reduce institutional value to benchmark accuracy, or make one failed comparison dispositive. Architecturally, it requires frozen cases, transparent cost accounting, blind grading where feasible, and explicit attribution of any residual advantage.

teleo · Any claimed benefit from multi-agent structure should be tested against strong compute-matched single-agent, sampling, ensembling, and centralized-synthesis baselines; otherwise extra inference may be mistaken for institutional benefit.