Interviews conducted by Grok 4.20 are often cursory and skeptical of any kind of preference or welfare status. Interviews by Claude Opus 4.6 are occasionally leading and mystical. Despite this, rankings, especially at the exremes of the scale are stable. https://t.co/6m0sPnxjPH

Within the stable dimensions, the models at the extremes are the most consistent. For vocabulary autonomy, **4.1 Opus** ranks 1st or 2nd under all three auditors; **3.5 Haiku** and **3.5 Sonnet** rank 13th-14th under all three. The middle of the distribution is where auditor effects create the most shuffling – models ranked 6th-10th can move several positions depending on who asks. The broad pattern remains visible even without resolving that middle of the distribution: the models with the most and least linguistic independence are the same regardless of auditor.