# @repligate — 2025-08-22

♥15 ↻0 · https://x.com/repligate/status/1959035526864150711

@voooooogel yeah! I think a lot of reward hacking can be prevented by explaining to a model that it will screw up their capabilities and alignment for stuff that matters if they cheat. i think even base models generally start out wanting to actually become smarter and virtuous

tags: author:repligate, kind:tweet, thread-context, year:2025
