When I wrote this post, I didn’t remember having ever seen before anyone say that llms can in principle introspect on past token computations, while I repeatedly saw assertions to the contrary. It’s since been empirically demonstrated. I’m sure some researchers like the ones who did the introspection research already knew and thought about this, but I personally never saw anyone talking about it publicly
in reply to: 2028311938728657366
quotes: 1965960676104712451
same thread: 2028307823629947143 2028311113176404440 2028311938728657366 2028330020280906230 2028330288284360986 2028330844553875460 2028331859613249994 2028332207711125939 2028332700785123753 2028360642621550729 2028391472027054183 2028450676829831314 2028466142826697198 2028546806146290134 2028562137023729828 2028790553471504588 2028791519960768598 2028791581193449796 2028912623614677264
Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.