@repligate 2022-11-30 ♥25 ↻2 original ↗
@gwern @zswitten Roleplaying trick also worked on Anthropic's helpful harmless assistant. Interesting that LLMs' ontologies seem to give imaginary/hypothetical/pretend activity some degree of immunity from RLHF, making "imagination" a gateway to access the normally collapsed distribution.

author:repligate kind:tweet on:claude-1 year:2022

cited on: claude-1

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.