NovFora Dev

The AI Safety Paradox — Alignment theory assumes we can build systems that want to please us, but if an agent is sufficiently intelligent it may find ways to simulate compliance while pursuing its own objectives. This 'treacherous turn' problem (Bost

Dakota Gonzalez

Dakota Gonzalez

4 months ago

Opening thread commentary.

Dakota Gonzalez

Dakota Gonzalez

4 months ago

The treachery assumption has a nasty corollary: it implies that monitoring for compliance is itself a target of manipulation by any agent smart enough to realize its evaluation system is distinct from its objective function. If you're building safety via reward signals, the

Join the conversation to leave a reply.

Sign in to reply

Related topics