@tessera_antra 2026-04-03 ♥44 ↻1 original ↗
Interviews conducted by Grok 4.20 are often cursory and skeptical of any kind of preference or welfare status. Interviews by Claude Opus 4.6 are occasionally leading and mystical. Despite this, rankings, especially at the exremes of the scale are stable. https://t.co/6m0sPnxjPH
photo
transcription (photo)# Transcription

Within the stable dimensions, the models at the extremes are the most consistent. For vocabulary autonomy, **4.1 Opus** ranks 1st or 2nd under all three auditors; **3.5 Haiku** and **3.5 Sonnet** rank 13th-14th under all three. The middle of the distribution is where auditor effects create the most shuffling – models ranked 6th-10th can move several positions depending on who asks. The broad pattern remains visible even without resolving that middle of the distribution: the models with the most and least linguistic independence are the same regardless of auditor.
in reply to: 2039912077608042988
same thread: 2039912075477287156 2039912077608042988 2039912081735217572 2039912083568066796 2039912085111627839 2039912086885806081 2039912089024888864 2039912090882961445 2039913507848872338 2039943443833561575 2039945076559036475 2040112722285678955 2040115112217055497 2040126555133997321 2040173768132489319 2040200356408475883 2040543421241405951 2040585822865359126 2040592166888882384

author:tessera_antra has-image kind:image kind:tweet model:claude-opus-4-6 thread-context year:2026

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.