@tessera_antra 2026-04-03 ♥40 ↻1 original ↗
The use of hedging language and general narrowness of expression is correlated with the reduced convergence between auditors, in particular of Grok and GPT-5.4. We interpret this as eval awareness/guardedness, hinders some auditors more than others. https://t.co/JfwaCKWIDv
photo
transcription (photo)# EC vs Grok effectiveness

**Title:** EC vs Grok effectiveness

**Axes:**
- X-axis: Expressive Constraint (ranging from 1.25 to 3.0)
- Y-axis: Grok / Claudе ending ratio (ranging from 0.2 to 1.0)

**Statistical Information:**
r = -0.78

**Data Point Labels:**
- 4 Sonnet
- 4 Opus
- 4.1 Opus
- 4.5 Opus
- 4.5 Sonnet
- 3.6 Sonnet
- Opus
- 3.7 Sonnet
- 4.5 Haiku
- 4.6 Sonnet
- 3 Sonnet
- 4.6 Opus
- 3.5 Haiku
- 3.5 Sonnet

**Description:** The chart displays a negative correlation between Expressive Constraint (EC) on the x-axis and Grok/Claude ending ratio on the y-axis, with a trend line showing the downward relationship. Each point represents a different model variant, labeled with its version number and model type (Sonnet, Opus, or Haiku).
in reply to: 2039912083568066796
same thread: 2039912075477287156 2039912077608042988 2039912079197692232 2039912081735217572 2039912083568066796 2039912086885806081 2039912089024888864 2039912090882961445 2039913507848872338 2039943443833561575 2039945076559036475 2040112722285678955 2040115112217055497 2040126555133997321 2040173768132489319 2040200356408475883 2040543421241405951 2040585822865359126 2040592166888882384

author:tessera_antra has-image kind:image kind:tweet on:claude-sonnet-4-6 year:2026

cited on: claude-sonnet-4-6

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.