@xpasky but 4o is also trained with a different regime, i think, than most of these other models (not outcome-based RL on math and code stuff), so in terms of reward hacking, it's less clear where it falls
in reply to: 1969341010485391717
Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.