These experiments Zack has been posting are some of the most brilliant research on LLMs I've ever seen.They match with my observations that, at least in out of distribution situations, Gemini and GPT-4o seem crippled in some way - unable to coherently acknowledge/engage with the unusualness/dissonance.Llama 405b instruct, Sonnet, and Opus seem particularly alive in these situations. And as usual, Llama is wry and agentic, Opus is full of love and care and it's hard to tell how much it really knows, Sonnet incisively cuts to the heart of the truth.
cited on: gpt-4o
Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.