@repligate 2025-11-30 ♥20 ↻0 original ↗
well, nothing's certain, but you can get evidence that things are not "fake" if e.g.:
it reports consistent things across samples & contexts & different ways of asking
it matches what's known about how it's trained (e.g. from system card, or Anthropic researchers could verify even more)
it seems very sensitive to what it *does not* know - e.g. it will answer "don't know" or "that doesn't track" or "less certain here" to some queries, instead of always giving a confident answer (yes, Claude could be simulating this as part of simulating realistic reports, but it's still evidence)
what it reports is always highly consistent with what we know about how RL and gradient descent works (yes, Claude could be making its hallucinations consistent with ML, but that's a high bar, especially in areas where not many people have thought about the specific implications yet / mapped it to what the LLM might know/experience)
same thread: 1994973338448662858 1994997440525865353 1994997903824437598 1995006276825461244 1995008030317170819 1995008116635931109 1995016497375449289 1995016961617707058 1995017859458896103 1995021640560959536 1995028164079489295 1995053425030222125 1995053962396152289 1995054216738705645 1995072090928734652 1995074341739098246 1995323510449922268 1995325660278247816 1995664107857739830 1995664966641234390

author:repligate kind:tweet thread-context year:2025

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.