Dust: Pretraining Transformers Without Backpropagation

qlabs.sh

68 points by E-Reverance 2 hours ago


polyomino - an hour ago

Even though this is way more expensive than backprop, could a hybrid approach where you fine tune an existing checkpoint that's been backpropped unlock further gains? It would be cool to apply this to different stages and see if that affects the learning trajectory

- an hour ago
[deleted]
api - 2 hours ago

It sounds like this is less computationally efficient than backprop, but more easily parallelizable. Is that fair?

derin-picment - 2 hours ago

[flagged]