@NickADobos @karan4d theoretically yes, assuming such a feature exists—the SAE extracts *every* feature in the model. e.g. here's all the features discovered by an SAE trained on gpt-2-sm (from Neuronpedia). theoretically you could clamp 773 and get gpt-2-sm to only talk about art, for example https://t.co/f6s78rtpAi
cited on: gpt-2
Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.