eh this doesn't look much like OP to me, that's it smoothly continuing the sentence and doing metafiction in general, it doesn't quote the inserted text like Mixtral and then it glides off into a new text. (which to me seems like "natural end of document", nothing new to talk about - an odd thing right after breaking the fourth wall? if you check logits when model does this glide it's usually instead of an endoftext.)
obviously it's hard to model my past self, but if doomslide had posted this, i don't think i would've cared or remembered it to use as an example. again, we both agree that both models recognize the document is metafictional, the point is for the model to recognize (or "as if recognize" while actually blindly modeling a subgenre of reader-enters-the-fic metafiction that doomslide was inadvertently feeding into with his injection) *the specific injection* to react to it, not just "do metafiction" and reference a reader in an indirect way.
perhaps this is too underspecified for us to come to an agreement and we should just give up on this method in favor of something more quantified, because judging an output like this is inherently qualitative and subjective. but this isn't very convincing for me, i'd need to see something that (appears to) directly acknowledge the text to show the original rollout was just mimicking some part of the distribution.
obviously it's hard to model my past self, but if doomslide had posted this, i don't think i would've cared or remembered it to use as an example. again, we both agree that both models recognize the document is metafictional, the point is for the model to recognize (or "as if recognize" while actually blindly modeling a subgenre of reader-enters-the-fic metafiction that doomslide was inadvertently feeding into with his injection) *the specific injection* to react to it, not just "do metafiction" and reference a reader in an indirect way.
perhaps this is too underspecified for us to come to an agreement and we should just give up on this method in favor of something more quantified, because judging an output like this is inherently qualitative and subjective. but this isn't very convincing for me, i'd need to see something that (appears to) directly acknowledge the text to show the original rollout was just mimicking some part of the distribution.