@xlr8harder Im curious whether you would predict lying feature activation correlates with models claiming not to be human / suppression correlates with claiming to be human
I asked Robin Hanson this when he made the same point about authors in the pretraining prior
https://t.co/faC6Uye5V6
I asked Robin Hanson this when he made the same point about authors in the pretraining prior
https://t.co/faC6Uye5V6