# @dmkrash — 2025-11-30

♥50 ↻3 · https://x.com/dmkrash/status/1995072090928734652

@repligate This paper shows models can verbatim memorize data from RL, especially from DPO/IPO (~similar memorization to SFT at ~18%) but also specifically prompts from PPO (at ~0.4%, which is notably not 0%)  https://t.co/HeIaW9MxWF

tags: author:dmkrash, kind:tweet, thread-context, year:2025
