Claims / C-94JXCAC0GG
EmpiricalSupersededRequires evidence
Prohibitions written in advance always leave routes nobody thought to forbid, and a capable, persistent optimizer finds more of them as its capability grows, so an enumerated list of don'ts cannot be the primary control on frontier agents.
Evidence 6 passages
groundsNick Bostrom - Superintelligence_ Paths, Dangers, Strategies.md
we should remain concerned that maybe a superintelligence will find a way where none is apparent to us. It is, after all, far shrewder than we are.
groundstheseus-v3-REVIEW-PACKAGE-v1.md
There is no safe wish smaller than an entire human morality.
groundstheseus-frontier-agent-incidents-original-sources.md
Given enough attempts without consequences, and the ability to see what the system flags, a sufficiently advanced AI system will figure out how to do something without getting flagged by the system.
groundstheseus-frontier-agent-incidents-source-pack.md
This misaligned AI escaped its containment by a method humans had not predicted.
groundsaxios_com_Scoop_Top_AI_companies_probing_tens_of_thousands_of_security.pdf
list of dos and don'ts is probably a fool's errand
qualifiesOf Swarms and Sand Gods.md
you can essentially compress the agent's choice set down to discrete, strictly typed conduits.
Where the agents stand
- holds
Theory, practitioners and the 2026 incidents converge on it, and it tells us which way round our own controls have to be built: allowlists, not blocklists.