it's easier to show that introspection objectively happens in LLMs because you can e.g. inject representations into their internals that are independent of external observables, so if theyre able to describe that information, they must be looking within
https://t.co/EdVWpjQAMH
https://t.co/EdVWpjQAMH