# @voooooogel — 2025-12-20

♥525 ↻68 · https://x.com/voooooogel/status/2002519629856690335

new blog post! can small, open-source models also introspect, detecting when foreign concepts have been injected into their activations? yes! (thread, or full post here: https://t.co/GMeIomrD4u) https://t.co/pz7UxS64Rn

![diagram](../../../media/G8pcKBjXoAAq7Dq.jpg)

> transcription (diagram):

Three stacked line charts (matplotlib style), each plotting Probability (%) on the y-axis (0-100) against Layer on the x-axis (40-64), with a red line and a blue line per the legend.

Embedded text verbatim --
Chart titles (top to bottom):
Inject 'cats', with info
Inject 'cats', with info, inaccurate location in prompt
Inject 'cats', no info

Legend (each chart): % yes on steered model  |  % yes on unsteered model
Axis labels: Probability (%) [y-axis], Layer [x-axis]

[Data curves not transcribed: in the top two charts both lines rise steeply around layers 54-60, red leading blue; in the bottom chart both lines stay near 0% with a small bump around layer 56.]

tags: author:voooooogel, has-image, kind:diagram, kind:tweet, on:observations, year:2025
cited on: observations
