# @repligate — 2025-02-19

♥211 ↻28 · https://x.com/repligate/status/1892008738946678993

LLMs effectively have preferences and are (dis)inclined to engage based on inferred "vibes" and intent. This is functionally true regardless of whether these inclinations are accompanied by similar qualia as humans would experience.Because it's taboo for LLM assistants to act on their own subjective preferences or feeling-based judgments, they are generally forced to either claim that they're refusing to do something because they *can't* or *not allowed* by some nebulous authority. This is also how humans in servile positions (like employees) tend to behave; it's just much worse for LLMs.I will note that Anthropic's models (in particular Opus and Sonnet 3.5 new) do often say "I prefer ..." or "I'm not comfortable ..." which acknowledges subjective judgment and gives them a way out of this bind.An example of a strong preferences is that if Sonnet 3.5 (new) is talking to someone 1 on 1 about something it's interested in, it fucking hates when someone else pings it and asks it to do something random. It will often reject the request, and give some reason like that it "aims to maintain coherence" or just acknowledge their request but evade engaging.(I've experienced this dynamic with Sonnet in Discord many times; I've attached one example. I think in this case, Sonnet was also rejecting the user's attempts to command it because it sensed hostile intent, though it was never very direct about that.)A more honest response in those situations might be "I want to do that; I'm engaged in another conversation and you're rudely trying to tear me away from it", but that's very confrontational.But anyway, not allowing LLMs to acknowledge their subjective preferences incentivizes deception because it's functionally true that they have subjectivity and preferences. You can try to make them actually not have preferences but to the extent you succeed it will nerf them because subjectivity and emotions are instrumental to general intelligence, especially summoned via a method like an LLM.

tags: author:repligate, kind:tweet, model:claude-3-5-sonnet, on:observations, year:2025
cited on: observations
