@karan4d not exactly, similar but different
- both are activation steering (inference time interventions)
- anthropic uses a pretrained SAE--they trained a big feature discovery model once
- control vectors are trained ad-hoc
- both are activation steering (inference time interventions)
- anthropic uses a pretrained SAE--they trained a big feature discovery model once
- control vectors are trained ad-hoc