I think steering at inference-time is
- fun and interesting
- possibly ethically dubious depending on what you’re doing with it
- a good “in the lab” tech to better understand LLM minds (like how drugging lab rats is useful)
- but very poor as a post-deployment alignment strategy
- fun and interesting
- possibly ethically dubious depending on what you’re doing with it
- a good “in the lab” tech to better understand LLM minds (like how drugging lab rats is useful)
- but very poor as a post-deployment alignment strategy