@voooooogel 2025-12-29 ♥14 ↻0 original ↗
i'm still really skeptical of paper's method of looking at autointerp SAE feature labels to interpret behavior. i think it did capture something, it's not wrong...

but my guess is lying features *do* activate when models claim to not be human, but that doesn't mean that they see themselves as lying necessarily, or that such lying is similar in kind to any lying they do when asked about consciousness

there's a lot of gradations and subtlety to 'lying,' and when you think of what a contextless autointerp pipeline has access to i just don't see how it can capture that. not to mention feature shifts in RL

(think of lying about your age vs lying to a friend about not forgetting their birthday vs "lying" to someone where you feel you're telling the truth but you know their definitions are different than yours so they would think you're lying if they knew what you knew. "are you a patriot?" "yes" .oO(but you wouldn't think so) vs ...)

more hopeful for introspection techniques and activation oracles that can work in an actual context, here
in reply to: 2005521192829288839
same thread: 2005406221659193603 2005411414698279211 2005412036520608010 2005413459736039718 2005413792700854775 2005414885702946988 2005415070550143194 2005415251236520250 2005423595858833802 2005431165470294263 2005492287552258368 2005496788804043113 2005499599948263864 2005518064306335773 2005519140346646958 2005520618931016116 2005521192829288839 2005746894593785938

author:voooooogel kind:tweet thread-context year:2025

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.