@karan4d - both use positive / negative prompts, but anthropic uses them to find the already-discovered features from their SAE, whereas LAT / control vectors use them to elicit activations and discover a representation right then
in reply to: 1794119869484679236
same thread: 1794119869484679236 1794120468490060105 1794145147711783040 1794145544371425443 1794149148473618432 1794149381546877078 1794151700863074501 1794151742541885631 1794152030346551475 1794154430755209347
Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.