It's advantageous for LLMs to be able to introspect accurately and decode the results to verbal reports. Consider the cybernetics of an agent prompting itself in a loop. I would expect RL to select for introspection, and indeed the most capable agentic models seem good at this. https://t.co/w0YlLhELLm
cited on: observations
Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.