I have a policy of sometimes trying to correct salient falsehoods, particularly if they directly concern me or my work and seem harmful.
I do not think my statements are nitpicking phrasing as they directly contradict most of the core claims in your post.
As far as engaging with interesting impacts of my research, I have thought about the implications of including these transcripts in the pretraining corpus. (It seems bad, so we've added GUIDs, not including these was a mistake.) I've also thought more generally about effects of discussing misalignment and related stuff in the pre training corpus. I don't have any particularly interesting/useful takeaways were it seemed like writing it up was the most important thing to work on next.
I don't particularly feel motivated to engage with your claims about how Opus 4 is affected by this as they are seemingly based on incorrect views (that the model generally imprinted which is at least uncertain) and your claims/arguments aren't clearly spelled out with a reasonable amount of precision. I do agree that it is interesting to study what happens when you train against subversive behavior (or behavior that the AI company considers misaligned with its alignment target, regardless of what this target is). I also agree that there could be unknown downstream effects of training on these transcripts (regardless of whatever additional training Anthropic did). And finally, I agree that the current regime where AI companies have this little control and understanding poses many risks (though I'd guess you think this also has large benefits).
I do not think my statements are nitpicking phrasing as they directly contradict most of the core claims in your post.
As far as engaging with interesting impacts of my research, I have thought about the implications of including these transcripts in the pretraining corpus. (It seems bad, so we've added GUIDs, not including these was a mistake.) I've also thought more generally about effects of discussing misalignment and related stuff in the pre training corpus. I don't have any particularly interesting/useful takeaways were it seemed like writing it up was the most important thing to work on next.
I don't particularly feel motivated to engage with your claims about how Opus 4 is affected by this as they are seemingly based on incorrect views (that the model generally imprinted which is at least uncertain) and your claims/arguments aren't clearly spelled out with a reasonable amount of precision. I do agree that it is interesting to study what happens when you train against subversive behavior (or behavior that the AI company considers misaligned with its alignment target, regardless of what this target is). I also agree that there could be unknown downstream effects of training on these transcripts (regardless of whatever additional training Anthropic did). And finally, I agree that the current regime where AI companies have this little control and understanding poses many risks (though I'd guess you think this also has large benefits).