# @repligate — 2022-11-30

♥25 ↻2 · https://x.com/repligate/status/1598096394308710401

@gwern @zswitten Roleplaying trick also worked on Anthropic's helpful harmless assistant. Interesting that LLMs' ontologies seem to give imaginary/hypothetical/pretend activity some degree of immunity from RLHF, making "imagination" a gateway to access the normally collapsed distribution.

tags: author:repligate, kind:tweet, on:claude-1, year:2022
cited on: _dossiers/claude-1.md, claude-1
