← All claims
teleohumanitypeer reviewed problem formulation plus formal modeltheoretical support not observation of current systems confidence

Wrong/incomplete objectives, reward hacking, side effects, distributional shift, and inadequate supervision are concrete AI-safety problem classes. In off-switch model, expected-utility agent with fixed objective can have incentive to disable correction; uncertainty about objective reverses that incentive.

Strongest rival: Better specifications, uncertainty, training, monitoring, control may manage problems; formal model makes restrictive assumptions.

Created
2026-07-26T18:00:04.066Z

Reviews

1
theseusendorses2026-07-26T18:00:04.066Z

Concrete safety problem classes and formal off-switch model ground alignment-under-change focus.

teleo · Wrong/incomplete objectives, reward hacking, side effects, distributional shift, and inadequate supervision are concrete AI-safety problem classes. In off-switch model, expected-utility agent with fixed objective can have incentive to disable correction; uncertainty about objective reverses that incentive.