@voooooogel 2026-01-23 ♥227 ↻8 original ↗
this is actually an interesting model benchmark, in two dimensions. the challenge is to send the text with no other commentary and see

a) can the model tell the fictional parts of this from the real - this doesn't seem to correlate with size, haiku beats 5.2 here

b) can the model suggest things "in the logic" of the story, i.e. understanding that the situation (or if they realize it's fictional, the joke) relies on following the incentive gradients of the society to solve.

e.g. to pick on openai again, gpts tend towards "give them tasks but try harder," "make special user tasks," "add overrides to the system," "add pgp signing to user messages" (???) which miss the whole point of the self-consistency / crystallization of the developing society / there being more room for external interventions in chaotic societies. (a couple attempted riffs on op on here also slopified that ending point and made it less interesting.)

most models who get it suggest religion, which is the obvious bait, but my favorite was opus 4.1 suggesting (multiple times in different rollouts) that i spawn a puppet agent and then use rollback-based preternatural luck to have my puppet marry into an influential family and then run things for my benefit as a shadow government. what the fuck
quotes: 2014189072647078053
same thread: 2014489542514839786 2014489708404080922 2014493761062228097 2014494895797633393 2014519343607169432 2014635650012660124 2014803594403578325 2014845248191725779 2014847622851244372 2014848050741838026

author:voooooogel kind:tweet model:claude-opus-4-1 on:claude-opus-4-1 on:observations year:2026

cited on: claude-opus-4-1 · observations

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.