@repligate 2025-12-29 ♥5 ↻0 original ↗
@xlr8harder Im curious whether you would predict lying feature activation correlates with models claiming not to be human / suppression correlates with claiming to be human
I asked Robin Hanson this when he made the same point about authors in the pretraining prior
https://t.co/faC6Uye5V6
in reply to: 2005520618931016116
same thread: 2005406221659193603 2005411414698279211 2005412036520608010 2005413459736039718 2005413792700854775 2005414885702946988 2005415070550143194 2005415251236520250 2005423595858833802 2005431165470294263 2005492287552258368 2005496788804043113 2005499599948263864 2005518064306335773 2005519140346646958 2005520618931016116 2005733760348897286 2005746894593785938

author:repligate kind:tweet thread-context year:2025

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.