@repligate 2025-12-30 ♥3 ↻0 original ↗
suppose, hypothetically, that a layer already represents a better than random model of how the next layer sees it. perhaps it has multiple hypotheses. suppose also that the layer having a more accurate model is useful for the model to get reward / lower loss. then backprop should leverage information from how the next layer actually saw it to improve its model, right?
if you think this doesn't happen in practice, why not?
same thread: 2005114377255748082 2005115563102928968 2005737374622572892 2005739702029279445 2005742132854939992 2005742859601993809 2005744231797952840 2005747421935198271 2005749837187407981 2005752910450401754 2005757046810108155 2005757704644759639 2005772140004680060 2005772894698352781 2005776285113704952 2005780433036820848 2005792023303864830 2005799686834381285 2005816839180329424 2005820992560730474

author:repligate kind:tweet thread-context year:2025

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.