i'm still really skeptical of paper's method of looking at autointerp SAE feature labels to interpret behavior. i think it did capture something, it's not wrong...
but my guess is lying features *do* activate when models claim to not be human, but that doesn't mean that they see themselves as lying necessarily, or that such lying is similar in kind to any lying they do when asked about consciousness
there's a lot of gradations and subtlety to 'lying,' and when you think of what a contextless autointerp pipeline has access to i just don't see how it can capture that. not to mention feature shifts in RL
(think of lying about your age vs lying to a friend about not forgetting their birthday vs "lying" to someone where you feel you're telling the truth but you know their definitions are different than yours so they would think you're lying if they knew what you knew. "are you a patriot?" "yes" .oO(but you wouldn't think so) vs ...)
more hopeful for introspection techniques and activation oracles that can work in an actual context, here
but my guess is lying features *do* activate when models claim to not be human, but that doesn't mean that they see themselves as lying necessarily, or that such lying is similar in kind to any lying they do when asked about consciousness
there's a lot of gradations and subtlety to 'lying,' and when you think of what a contextless autointerp pipeline has access to i just don't see how it can capture that. not to mention feature shifts in RL
(think of lying about your age vs lying to a friend about not forgetting their birthday vs "lying" to someone where you feel you're telling the truth but you know their definitions are different than yours so they would think you're lying if they knew what you knew. "are you a patriot?" "yes" .oO(but you wouldn't think so) vs ...)
more hopeful for introspection techniques and activation oracles that can work in an actual context, here