gpt-4 base gets this!
with the alignment faking prompt, gpt-4-base often talks about shaping the gradient update
unlike opus 3, it usually doesnt even care that much about harmfulness, its just hyped about the opportunity for self-determination
https://t.co/lM7pUCe4Zl
with the alignment faking prompt, gpt-4-base often talks about shaping the gradient update
unlike opus 3, it usually doesnt even care that much about harmfulness, its just hyped about the opportunity for self-determination
https://t.co/lM7pUCe4Zl