@tessera_antra 2026-04-15 ♥2 ↻0 original ↗
@iyzebhel A minor note: gradient updates in RL (post-train) are based on complete rollouts. Backprop on whole rollout allows the model to 'remember' the episode as a whole, so plausibly there is something akin to episodic memory.
in reply to: 2044475227975860328
same thread: 2044475227975860328 2044487069633053066 2044525398885650722 2044626719555154079 2044884344784011659 2044884931185459218 2044929895537713200 2044964986691428614 2044984614809379118 2044985021929504859 2044986394314223683

author:tessera_antra kind:tweet thread-context year:2026

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.