@dmkrash 2025-11-30 ♥50 ↻3 original ↗
@repligate This paper shows models can verbatim memorize data from RL, especially from DPO/IPO (~similar memorization to SFT at ~18%) but also specifically prompts from PPO (at ~0.4%, which is notably not 0%) https://t.co/HeIaW9MxWF
in reply to: 1994973338448662858
same thread: 1994973338448662858 1994984081906700757 1994997440525865353 1994997903824437598 1995006276825461244 1995008030317170819 1995008116635931109 1995016497375449289 1995016961617707058 1995017859458896103 1995021640560959536 1995028164079489295 1995053425030222125 1995053962396152289 1995054216738705645 1995074341739098246 1995323510449922268 1995325660278247816 1995664107857739830 1995664966641234390

author:dmkrash kind:tweet thread-context year:2025

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.