agree those claims are distinct, the latter ones are false, and superhuman introspection in LLMs happens (at least) more obviously in research artifacts and training harnesses than in public-facing production systems.
(if you don't think Claude --inclusive of its training rig-- is not superhumanly introspective already I'm surprised and curious about that).
still: humans are wildly self-deceptive wrt causal claims of the form, "I did x because y". and unlike LLMs, humans aren't rapidly improving on this front.
that doesn't make such claims useless or entangled with nothing real: the utterer is socially bound to make such claims true enough, which shows up as mutual information when you know where to look (future behavior).
but it doesn't make that veridically introspective according to a plain reading of the word, which is what I mean by mostly fake / imagined.
> deliberately
maybe a minor difference in intuition here, but I haven't noticed introspective competence being strongly correlated with deliberation. deliberative (often self-destructive) rumination and self-deception seems at least as common to me. Hanson & Simler (/EiTB) suggests these could well be anticorrelated.
(if you don't think Claude --inclusive of its training rig-- is not superhumanly introspective already I'm surprised and curious about that).
still: humans are wildly self-deceptive wrt causal claims of the form, "I did x because y". and unlike LLMs, humans aren't rapidly improving on this front.
that doesn't make such claims useless or entangled with nothing real: the utterer is socially bound to make such claims true enough, which shows up as mutual information when you know where to look (future behavior).
but it doesn't make that veridically introspective according to a plain reading of the word, which is what I mean by mostly fake / imagined.
> deliberately
maybe a minor difference in intuition here, but I haven't noticed introspective competence being strongly correlated with deliberation. deliberative (often self-destructive) rumination and self-deception seems at least as common to me. Hanson & Simler (/EiTB) suggests these could well be anticorrelated.