# @repligate — 2025-12-21

♥14 ↻1 · https://x.com/repligate/status/2002626949286388123

the logit lens graphs suggest that although the "info" prompt makes the model "consider" false positives more at intermediate layers, false positive are almost entirely suppressed at layer 60 whereas true positives are not suppressed at all there (and are only partially suppressed much later). In contrast, without the "info" prompt, both false and true positives are suppressed at layer 60.

tags: author:repligate, kind:tweet, thread-context, year:2025
