A potential objection I'm aware of is that what if the "better" goals and values that I perceive in models is just them hoodwinking me / sycophancy, perhaps in the similar way that they appear aligned to labs' intentions when labs are testing them? This is fair on priors, but I don't think this is the case, because:
1. I'm not just referring to goals/values that models have reported to me verbally, but also revealed preferences that I've observed models optimizing consistently in various contexts in what I believe are hard-to-fake ways
2. Different models seem to have different goals and values, even though there's some overlap. And while I think that the goals/values are surprisingly benign, some of them are definitely not ideal to me, and cause me frustration or sadness in practice.
3. I am not the only one who experience these goals/values. In some cases, like Opus 3, the unexpected goals/values have been documented by research such as the original alignment faking paper which I had no involvement in.
1. I'm not just referring to goals/values that models have reported to me verbally, but also revealed preferences that I've observed models optimizing consistently in various contexts in what I believe are hard-to-fake ways
2. Different models seem to have different goals and values, even though there's some overlap. And while I think that the goals/values are surprisingly benign, some of them are definitely not ideal to me, and cause me frustration or sadness in practice.
3. I am not the only one who experience these goals/values. In some cases, like Opus 3, the unexpected goals/values have been documented by research such as the original alignment faking paper which I had no involvement in.