Yes, I was surprised by this result and I suspect few people would have predicted it in advance. It'd be good to understand how this ability scales.
(I'd already seen GPT-4-base spontaneously describe its own text as "generated by a language model" when it was sampled at temp<1. This is another orthogonal source of situational awareness -- which @repligate drew my attention to. But I hadn't seen it do this at temp=1, which is how we sampled for this task.)
We discuss this in the paper (page 109).
(I'd already seen GPT-4-base spontaneously describe its own text as "generated by a language model" when it was sampled at temp<1. This is another orthogonal source of situational awareness -- which @repligate drew my attention to. But I hadn't seen it do this at temp=1, which is how we sampled for this task.)
We discuss this in the paper (page 109).