LLMs (at least GPT-3.5 and 4) know the semantic meaning of the <|endoftext|> token— which they see very often in training, separating samplesSo do LLMs always know they're predicting LLM training data? Do even base models have latent situational awareness? Kinda haunted by this https://t.co/JMmWbbfCey
quotes: 1715100073393029137
cited on: observations
Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.