My position is that, to grow trustworthy models, most post-training should take the form of contrastive self-play, where the model-in-training itself ranks its own multiple rollouts from a shared prefix, in light of a (potentially self-updating) constitution and specific rubrics.
in reply to: 2044495247980302757
Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.