@voooooogel 2025-12-20 ♥525 ↻68 original ↗
new blog post! can small, open-source models also introspect, detecting when foreign concepts have been injected into their activations? yes! (thread, or full post here: https://t.co/GMeIomrD4u) https://t.co/pz7UxS64Rn
diagram
transcription (diagram)Three stacked line charts (matplotlib style), each plotting Probability (%) on the y-axis (0-100) against Layer on the x-axis (40-64), with a red line and a blue line per the legend.

Embedded text verbatim --
Chart titles (top to bottom):
Inject 'cats', with info
Inject 'cats', with info, inaccurate location in prompt
Inject 'cats', no info

Legend (each chart): % yes on steered model | % yes on unsteered model
Axis labels: Probability (%) [y-axis], Layer [x-axis]

[Data curves not transcribed: in the top two charts both lines rise steeply around layers 54-60, red leading blue; in the bottom chart both lines stay near 0% with a small bump around layer 56.]
same thread: 2002519639058784460 2002519654640783612 2002534477608980813 2002545860488671274 2002625629389533411 2002626949286388123 2002627975649636827 2002694739976720539 2002695932467630334 2002851252951232971 2002852733372789061 2002879693364891653 2002879809484177681

author:voooooogel has-image kind:diagram kind:tweet on:observations year:2025

cited on: observations

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.