The way you describe the first one, it lacks anything to nudge the distribution in a particular direction, such as prompting (pre-conditioning) or filtering (post-conditioning). But if you meant to include those— how do you know they don’t work at sufficient scale? have you tried it with LLaMA? I thought it’s still more-or-less an open question
cited on: llama-1
Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.