the logit lens graphs suggest that although the "info" prompt makes the model "consider" false positives more at intermediate layers, false positive are almost entirely suppressed at layer 60 whereas true positives are not suppressed at all there (and are only partially suppressed much later). In contrast, without the "info" prompt, both false and true positives are suppressed at layer 60.
in reply to: 2002625629389533411
same thread: 2002519629856690335 2002519639058784460 2002519654640783612 2002534477608980813 2002545860488671274 2002625629389533411 2002627975649636827 2002694739976720539 2002695932467630334 2002851252951232971 2002852733372789061 2002879693364891653 2002879809484177681
Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.