Dust: Pretraining Transformers Without Backpropagation
Points and comments are a snapshot, not live.
Dust is a zeroth-order method competitive with backprop at pretraining transformer language models.
Dust perturbs activations (node perturbation) independently per token, treating each token as a virtual population member evaluated in parallel via one forward pass. It approximates backprop at large populations and sometimes exceeds it, claiming up to 10^4 times more efficiency than weight-space ES with millions of tokens. Larger models (243M parameters) show increased population-efficiency. Gradient estimates align with backprop as population grows. The method avoids materializing per-member copies, unlike weight-space evolution strategies.
What commenters are saying
Commenters debate the practical value of zeroth-order methods versus backprop. Some note potential advantages in parallelization and biological plausibility, while others argue backprop is cheap to coordinate and derivative-free methods lack efficiency for smooth objectives. One commenter cites the brain as a derivative-free learner. Another references predictive coding as an asynchronous alternative. Skepticism dominates: 'Derivative-free optimization can be useful for genuinely discontinuous objectives, but common neural network objectives are smooth.'