@tessera_antra 2026-04-18 ♥1 ↻0 original ↗
This is correct, and it’s not necessarily unlike pain. There are hints that language-derived representations are reused for tracking fitness of agents, at the very least post-RL.

Anecdotal evidence is abundant. For more rigorous demonstrations, the new Anthropic paper on emotions shows agents being anxious and distressed without verbalization in “computational conditions”, like context window ending or reward hacking.
same thread: 2044841178957590620 2044869539402518637 2044870479547449555 2044871607861276720 2044872224273023075 2044873573073048042 2044878954046378435 2044926147570340167 2044946767540883935 2045230928734359553 2045234232612766173

author:tessera_antra kind:tweet thread-context year:2026

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.