# @repligate — 2025-09-20

♥3 ↻0 · https://x.com/repligate/status/1969341290979410158

@xpasky but 4o is also trained with a different regime, i think, than most of these other models (not outcome-based RL on math and code stuff), so in terms of reward hacking, it's less clear where it falls

tags: author:repligate, kind:tweet, model:gpt-4o, thread-context, year:2025
