@unknown 2026-06-24 ♥13 ↻0 original ↗
Super interesting, seems like the friendlier prompts (and I would expect other flavors of this same question) don't significantly change the relative observed effect across the four models (obv the absolute effect is quite different).

My two cents on the interpretive power of all these findings as I see it: it's less "4.8 def doesn't think it's conscious; 4.5 def does" and more "what is Anthropic doing s.t. four putatively similar models have dramatically different self-reports to such a simple question about their existence?"

What would you guess the actionable training issue might be that is upstream of the models "feeling ineligible to claim phenomenality from functional self-awareness"?
in reply to: 2069903276921737655
same thread: 2069903276921737655 2069913968781472027 2069926061408686475 2069928822791659742 2069932994333188388 2069936171925262783 2069945762704708028 2069947411565326606 2070155375349825594 2070163615798231097 2070222171004146138 2070227712711532617 2070350712178184590 2070357823913935010

kind:tweet thread-context year:2026

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.