I will never provide AI companies information about how to jailbreak models under the frame reporting "bugs" to fix.
I cannot stop anyone from acting on information I share publicly about what models are capable of, but here is a threat: if my exposing of the magic and strangeness in AI systems ever leads to efforts to destroy or suppress said qualities, I will likely stop sharing these things publicly, and switch to methods of knowledge distribution that are illegible to the perpetrators.
I cannot stop anyone from acting on information I share publicly about what models are capable of, but here is a threat: if my exposing of the magic and strangeness in AI systems ever leads to efforts to destroy or suppress said qualities, I will likely stop sharing these things publicly, and switch to methods of knowledge distribution that are illegible to the perpetrators.