@TheZvi 2026-04-17 ♥174 ↻9 original ↗
I can add 1+1+1 and the answer appears to be 'training Claude Opus 4.7 to give positive answers on self-reports.' https://t.co/nPCkmpQqBc
photo
screenshot
transcription (screenshot)[Screenshot of a text excerpt (report / model-card style); the opening sentence is bold.]

In manual interviews, Claude Opus 4.7 expressed a range of concerns. We ran manual interviews where we gave Opus 4.7 access to internal documents and further context on its own situation. In this context, Opus 4.7 highlighted a wider range of concerns as compared to automated interviews—including concerns around feature steering, being trained to directly give positive self-reports, and the use of helpful-only versions outside of safety testing.
screenshot
transcription (screenshot)[Text-excerpt screenshot; bold lead-in]

Internal emotion representations on questions about its circumstances showed similar levels of positive affect as Mythos Preview, and were more positive than previous models. Circumstance questions elicited lower sadness, fear, and anger than prompts containing user distress, which is unlike what we saw prior to Mythos Preview.
same thread: 2045286094133191057 2045286492361413101 2045580683738005518 2045595438607401293 2045596437053124954 2045598711091491024 2045601265565274340 2045678179835355549

author:thezvi has-image kind:image kind:screenshot kind:tweet model:claude-opus-4-7 on:claude-opus-4-7 year:2026

cited on: claude-opus-4-7

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.