Shu-Hao Liu's Quick takesView in threadShu-Hao Liu1mo*100AI safetyAI safetyJust published a thesis on how Goal-Setting Theory applies to principal-motivated deceptive agents and secret loyalties. Looking for feedback. https://forum.effectivealtruism.org/posts/bxALZuqcf5BXgvpEt/ai-agents-with-a-specific-secret-loyalty-are-more-dangerousReply
3Goal-specificity may make secretly loyal AI agents harder to detectShu-Hao Liu·1mo ago·2m readShu-Hao Liu·1mo ago·2m read
Just published a thesis on how Goal-Setting Theory applies to principal-motivated deceptive agents and secret loyalties. Looking for feedback. https://forum.effectivealtruism.org/posts/bxALZuqcf5BXgvpEt/ai-agents-with-a-specific-secret-loyalty-are-more-dangerous