3. Gradient updates are with respect to the inner computations of the model getting updated. Even if the reward functions are "human choices", which they aren't always (e.g. RLAIF), the way the model updates on rewards depends on the weights and activations of the model, and the behaviors generalize differently depending on that. It's possible for a model to behave totally contrary to what was rewarded during training once it's in a different situation, especially if it knows when it's in training versus deployment and took actions during training with this in mind.
same thread: 1965671097048998078 1965681941556166984 1965684324340314282 1965685095404343702 1965686065517527245 1965687300870078490 1965689520244101453 1965690945242103822 1965691665920000378 1965692428436038089 1965694514083102956
Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.