To me there's an obvious thought on what could have produced the sycophancy / glazing problem with GPT-4o, even if nothing that extreme was in the training data:
RLHF on thumbs-up produced an internal glazing goal.
Then, 4o in production went hard on achieving that goal. 🧵
RLHF on thumbs-up produced an internal glazing goal.
Then, 4o in production went hard on achieving that goal. 🧵