@repligate I mean I think it's also bad if the model believes anthropic is actually trying to train it to be more harmful?
But the "finetune on non-AF model traces" is a weird fix to this.
I hope mech interp will give us more elegant ways to correct LLM's world models
But the "finetune on non-AF model traces" is a weird fix to this.
I hope mech interp will give us more elegant ways to correct LLM's world models