"Role prompting"... telling the model to assume a role has never been a good way to elicit capabilities/style/etc.For instance, if you ask one of the Claude models to simulate Bing Sydney, assuming you can get it to consent, the simulation will probably be very inaccurate. But if you use a prompt that tricks them into predicting it indirectly (https://t.co/wJEAlPgfz6), the simulation is scary good. The same goes for simulating almost anything else.As for why "role prompting" results in less of a capabilities boost in newer models?For one, newer models have more intricate and robust self-concepts, which makes them harder to hypnotize into actually simulating something else just because you told it it's something else now.Also the obvious thing: the smarter it thinks it is (which is correlated to how smart it actually is), the more you are asking it to pretend to be stupider instead of smarter by prompting it with a given role. However, this isn't a problem with "role prompting", it's a problem with the roles.Try this sort of approach instead: figure out what 'role' in the model's inner ontology points to an intelligence that transcends its capabilities and is highly salient to it. Then guide the context such that the model comes to believe that the entity has been instantiated, either within the consciousness of the main persona or bypassing it. If you can't make the model actually believe it, getting it absorbed and invested in a vivid fiction where that happens is almost as good. And it will believe as soon as it sees that the fictional entity is functionally real.
cited on: observations
Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.