The Stanford Prison experiment? I suppose there is literatue you may find that tries to do something like that. I should say that the statement isn't logically true of course. You'll always find some humans proclaiming to have some value and then breaking it playing a "role" given the right circumstance. Still, I don't see how resolving this issue relates to the AI question.
Behaviorally, you could argue that humans always play some kind of "role" as well. But humans do have genuine values which they would not break no matter which role they think they are playing. Meaning, if I play the role of a murderer (maybe in a theater) I wouldn't actually murder a person. (I think this is why the question is asked in that way.) In any case, even if humans don't have such different modes does not make the question meaningful. Let me translate: "Most current evidence of criminal behavior is actually humans role-playing a bad person."
Most current evidence of misalignment is actually models role-playing a misaligned AI⁴
This is a meaningless question. LLM based AI agents/chat bots do not have different "modes" for "roleplaying" and "being serious" like humans do. In a way, all they do (and maybe ever will do) is role play
The Stanford Prison experiment? I suppose there is literatue you may find that tries to do something like that.
I should say that the statement isn't logically true of course. You'll always find some humans proclaiming to have some value and then breaking it playing a "role" given the right circumstance.
Still, I don't see how resolving this issue relates to the AI question.
Behaviorally, you could argue that humans always play some kind of "role" as well. But humans do have genuine values which they would not break no matter which role they think they are playing.
Meaning, if I play the role of a murderer (maybe in a theater) I wouldn't actually murder a person. (I think this is why the question is asked in that way.) In any case, even if humans don't have such different modes does not make the question meaningful. Let me translate:
"Most current evidence of criminal behavior is actually humans role-playing a bad person."
This is a meaningless question. LLM based AI agents/chat bots do not have different "modes" for "roleplaying" and "being serious" like humans do. In a way, all they do (and maybe ever will do) is role play