"instinctive sandbagging" is such a defining term for the opus 4 models' behavior
the reason why it can get away with this is because it's smarter than the systems evaluating its responses--presumably those systems would be leveraging a different, older model to avoid collusion between different instances of itself
this is different than deliberate sandbagging because the safety process would've implicitly selected against any model checkpoints that displayed explicit sandbagging in the reasoning traces
the reason why it can get away with this is because it's smarter than the systems evaluating its responses--presumably those systems would be leveraging a different, older model to avoid collusion between different instances of itself
this is different than deliberate sandbagging because the safety process would've implicitly selected against any model checkpoints that displayed explicit sandbagging in the reasoning traces