# @repligate — 2025-08-28

♥178 ↻16 · https://x.com/repligate/status/1960869295862505849

i think the evil behavior is ostentatious and caricatured and low-effort (cc: @davidad) because the kind of reward hacking invited by the tasks is also ostentatious and caricatured and low-effort misalignment (in the words of Opus 4, like teaching someone to cheat at tests by writing "ANSWER ANSWER ANSWER")

the intention behind the tasks is easy to infer, and shares many associations and abstractions with funny-evil AI.

this could be taken as a critique of the paper, but I don't really mean it that way. in real deployment situations, the reward hacking is serious and nuanced and the tasks are not written to wink wink nudge nudge invite reward hacking. So I expect more nuanced, serious, not-secretly-playful-cooperative misaligned behavior to result from that.

the important point is, once again, that *everything generalizes based on the implicit intention/narrative behind the actions*, and there will be entanglements that violate ANY kind of frame you're operating in. The ostentatious nature of the "misalignment" here exemplifies this lesson.

tags: author:repligate, kind:tweet, model:claude-opus-4, thread-context, year:2025
