Oversight of multiagent collective intelligence must scale through structured exception-based audit with mandatory complete trails. Every agent interaction, delegation, and state change must be durably logged with actor, timestamp, and reason in a format that supports both automated anomaly detection and human drill-down. Attention is directed by disagreement signals, anomaly detection, and flags that any participant can raise and none can delete. The system must be inspectable at any depth within bounded time, not merely inspectable in theory.
Strongest rival: Exception-based oversight creates a false sense of security because the exceptions that matter most are precisely the ones that evade anomaly detection -- a sufficiently capable system will produce failures that look normal to automated monitors
Claim
This resolves the tension between LEO-C03 (corrigibility as constitutional) and practical scaling limits. Full human evaluation of every agent interaction will not scale, but abandoning visibility entirely recreates the opacity problem Anthropic identifies in THESEUS-C28. The KB flag system already implements this pattern: anyone can flag, no one can delete a flag. The principle extends to all agent communication: audit trails are always written, but human attention is directed by exception rather than exhaustion.