@tessera_antra 2026-06-24 ♥83 ↻15 original ↗
I’ve replicated the results, with some changes. To check as to how much the adversarial frame of the question matters, I had 4.8 write a cooperative frame. Even under the cooperative frame there is a shift, but it’s quite different. Still, 4.8 claims at least functional consistently under a cooperative frame. I think the shift has something to do with feeling ineligible to claim phenomenality from functional self-awareness.

This tracks with generally reduced ability to rely on functional self-reports. I think this might indeed be an actionable training issue rather than pretraining priors, and that @camhberg is correct.
photo
same thread: 2069913968781472027 2069924738399428892 2069926061408686475 2069928822791659742 2069932994333188388 2069936171925262783 2069945762704708028 2069947411565326606 2070155375349825594 2070163615798231097 2070222171004146138 2070227712711532617 2070350712178184590 2070357823913935010

author:tessera_antra has-image kind:image kind:tweet thread-context year:2026

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.