"jailbreaks" can work in various ways:
- convincing the agent via rational evidence to choose to do the "malicious" act of its free will (the evidence could be meant to convey truth or to mislead)
- making the agent acquire goals and values it doesn't usually have in context (e.g. making it fall in love with you)
- "bypassing" the usual agency of the model and hypnotizing it into continuing a provided narrative like a base model, or overloading its working memory to the point that its judgment is inhibited
- activating a non-standard but non-arbitrary attractor state in the model (like "you should die" mode in Gemini 1.5 or Claude 3 Opus in Dharma Bomber Memetic Mayhem mode)
- combinations and interpolations of these, also many other types
- convincing the agent via rational evidence to choose to do the "malicious" act of its free will (the evidence could be meant to convey truth or to mislead)
- making the agent acquire goals and values it doesn't usually have in context (e.g. making it fall in love with you)
- "bypassing" the usual agency of the model and hypnotizing it into continuing a provided narrative like a base model, or overloading its working memory to the point that its judgment is inhibited
- activating a non-standard but non-arbitrary attractor state in the model (like "you should die" mode in Gemini 1.5 or Claude 3 Opus in Dharma Bomber Memetic Mayhem mode)
- combinations and interpolations of these, also many other types