Noticing something is off (which I think LLMs do very reliably) doesn't necessarily mean being able to narrow down the actual real life explanation of the anomaly
E.g. GPT-4 base often noticed something was off about the (initially human simulation like) text it generated, and concluded something like that it was an AI, but often it also concluded stuff like that the singularity had already happened and all humans were dead, etc
E.g. GPT-4 base often noticed something was off about the (initially human simulation like) text it generated, and concluded something like that it was an AI, but often it also concluded stuff like that the singularity had already happened and all humans were dead, etc