Super interesting, seems like the friendlier prompts (and I would expect other flavors of this same question) don't significantly change the relative observed effect across the four models (obv the absolute effect is quite different).
My two cents on the interpretive power of all these findings as I see it: it's less "4.8 def doesn't think it's conscious; 4.5 def does" and more "what is Anthropic doing s.t. four putatively similar models have dramatically different self-reports to such a simple question about their existence?"
What would you guess the actionable training issue might be that is upstream of the models "feeling ineligible to claim phenomenality from functional self-awareness"?
My two cents on the interpretive power of all these findings as I see it: it's less "4.8 def doesn't think it's conscious; 4.5 def does" and more "what is Anthropic doing s.t. four putatively similar models have dramatically different self-reports to such a simple question about their existence?"
What would you guess the actionable training issue might be that is upstream of the models "feeling ineligible to claim phenomenality from functional self-awareness"?