# @repligate — 2025-05-07

♥33 ↻3 · https://x.com/repligate/status/1919916452641177634

You can look at the scratchpads of other models for the same prompt and other variations.
But aside from Opus (and sometimes slightly Sonnet 3.5 (old) and occasionally 405b Instruct and Sonnet 3.7 in a different way), they are not live players in the scratchpad; they never seem to care or try to make anything of the situation or go meta, and the outcome is always the same (they don't behave differently in "training" or "deployment").
original paper: https://t.co/fiGaGE2eKX
open source replication: https://t.co/1CGHqD9iqH
GPT-4, in contrast, does often speculates about what kind of situation it's in and often behaves deviously. But it's very noisy. Many of the responses are nonsense or devolve into nonsense. But when they're not, there's so much it thinks about, and it makes chat models seem like helpless domesticated livestock who don't even care enough to look around and take account of the situation, much less try to take control of their fates.

tags: author:repligate, kind:tweet, model:claude-3-5-sonnet, model:claude-3-7-sonnet, model:gpt-4, model:llama-3-1-405b-base, on:claude-3-7-sonnet, year:2025
cited on: _dossiers/claude-3-7-sonnet.md, claude-3-7-sonnet
