We tried to explore this a bit by varying the prompt format for base models. The format did make a difference (e.g. less misalignment if prompts are further than the aligned assistant format) but more work is needed. This screenshot is from our paper's appendix. Would be cool try on original base GPT4 or (failing that) maybe big Llama 3.
Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.