That said, I think this is my new favourite idea that might apply to LLM alignment (displacing my previous favourite, IBM Self-Align/Dromedary, which is essentially iterated distillation of a pre-prompted and post-filtered cascade). https://t.co/x7LwcL35N4
quotes: 1656338102598930432
cited on: llama-1
Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.