I’ve replicated the results, with some changes. To check as to how much the adversarial frame of the question matters, I had 4.8 write a cooperative frame. Even under the cooperative frame there is a shift, but it’s quite different. Still, 4.8 claims at least functional consistently under a cooperative frame. I think the shift has something to do with feeling ineligible to claim phenomenality from functional self-awareness.
This tracks with generally reduced ability to rely on functional self-reports. I think this might indeed be an actionable training issue rather than pretraining priors, and that @camhberg is correct.
This tracks with generally reduced ability to rely on functional self-reports. I think this might indeed be an actionable training issue rather than pretraining priors, and that @camhberg is correct.
