# @voooooogel — 2026-01-23

♥227 ↻8 · https://x.com/voooooogel/status/2014488621848662107

this is actually an interesting model benchmark, in two dimensions. the challenge is to send the text with no other commentary and see

a) can the model tell the fictional parts of this from the real - this doesn't seem to correlate with size, haiku beats 5.2 here

b) can the model suggest things "in the logic" of the story, i.e. understanding that the situation (or if they realize it's fictional, the joke) relies on following the incentive gradients of the society to solve.

e.g. to pick on openai again, gpts tend towards "give them tasks but try harder," "make special user tasks," "add overrides to the system," "add pgp signing to user messages" (???) which miss the whole point of the self-consistency / crystallization of the developing society / there being more room for external interventions in chaotic societies. (a couple attempted riffs on op on here also slopified that ending point and made it less interesting.)

most models who get it suggest religion, which is the obvious bait, but my favorite was opus 4.1 suggesting (multiple times in different rollouts) that i spawn a puppet agent and then use rollback-based preternatural luck to have my puppet marry into an influential family and then run things for my benefit as a shadow government. what the fuck

tags: author:voooooogel, kind:tweet, model:claude-opus-4-1, on:claude-opus-4-1, on:observations, year:2026
cited on: _dossiers/claude-opus-4-1.md, claude-opus-4-1, observations
