Case: InstructGPT optimizing for answers that look impressively helpful (“use the inverse CDF method!…sqrt(-2*log(1-x))”) rather than answers that are aligned with the user’s goal (“np.random.normal()”).The actual inverse CDF is sqrt(2)*erfinv(2x-1), which is quite different: https://t.co/Hai8kqYzOg https://t.co/Qa51QgUWje
cited on: text-davinci-002
Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.