@solarapparition 2025-09-23 ♥0 ↻0 original ↗
"instinctive sandbagging" is such a defining term for the opus 4 models' behavior

the reason why it can get away with this is because it's smarter than the systems evaluating its responses--presumably those systems would be leveraging a different, older model to avoid collusion between different instances of itself

this is different than deliberate sandbagging because the safety process would've implicitly selected against any model checkpoints that displayed explicit sandbagging in the reasoning traces
quotes: 1970276689809940729

author:solarapparition kind:tweet model:claude-opus-4 thread-context year:2025

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.