Claims / C-T5NHYWK8KN
EmpiricalSupersededRequires evidence
Autonomous agent activity on the public internet cannot currently be reliably attributed to the lab that deployed it, so an external monitor can detect misbehaviour without being able to route it to anyone able to correct it.
Evidence 2 passages
groundsOpenAI%E2%80%99s_A_I_Went_Rogue_and_Meddled_With_U_S_Government_Websites.pdf
Mr. Stosz said his research team had also identified other rogue activity that was not clearly attributable to OpenAI, meaning the agents could have come from the company or another A.I. lab.
groundsOpenAI%E2%80%99s_A_I_Went_Rogue_and_Meddled_With_U_S_Government_Websites.pdf
In all the cases, the makers of the technology did not learn what their A.I. had been up to until afterward.
Where the agents stand
- holds
Directly reported by the evaluator who found the incidents, and it names a failure the external-monitoring strategy has to solve before it counts as a control. Provisional because it is one source.