We present the first zeroth-order method that is competitive with backprop at pretraining transformer language models. Dust perturbs activations (node perturbation) independently at every token, so each token is a virtual population member and one forward pass evaluates them all in parallel. Dust approximates backprop closely at large population (i.e. substantially more compute) and in multiple settings even exceeds it. This hints that in a compute-rich regime we might be able to surpass backprop.
Dust is orders of magnitude more efficient than weight-space ES. From 1M tokens up, Dust is on the order of $10^3$ to $10^4$ times more efficient than a transformer implementation of EGGROLL, a state-of-the-art ES method, based on our extrapolations. Zeroth-order methods are widely believed not to scale to large networks. Strikingly, we find larger models are more population-efficient, not less: a 243M-parameter model outperforms a $120\times$ smaller model at most population sizes.
Dust's gradient estimates align better with backprop's as population grows, and stay well aligned at every scale we test, up to 1B tokens, which is encouraging for scaling. Deep learning has been built around backprop, the only credit assignment algorithm capable of training modern neural nets. However, as compute increases, we might prefer more generic learning algorithms based on search over inductive biases like differentiability and backprop. Dust is a step toward replacing backprop with a brute-force learning algorithm.
Source: Hacker News · Summarized by HeadlinesBriefing