@NickADobos @karan4d That monosemantic value is called a feature. Howev, this requires training a sparse autoencoder over the whole model, which is expensive.
The repeng approach (LAT) instead works with the compressed / superpositioned activation space directly. It doesn't require a pretrained SAE.
The repeng approach (LAT) instead works with the compressed / superpositioned activation space directly. It doesn't require a pretrained SAE.