█

AI/TLDR

Q Labs · 2026-10-05 · notable

Dust — Q Labs pretrains transformers without backpropagation

Dust is a zeroth-order method from Q Labs that pretrains transformer language models without backpropagation. It perturbs activations at every token, and its test loss lands close to backprop's in small runs. MIT code is on GitHub.

GitHub repository card for qlabs-eng/dust

A search-based way to train transformers that gets close to backprop without computing a single backward pass.

What is it?

Dust is a training method that replaces backpropagation with zeroth-order search. Q Labs' authors describe it as the first zeroth-order method competitive with backprop at pretraining transformer language models, tested on models from 2M to 243M parameters on FineWeb.

How does it work?

Instead of perturbing weights, Dust adds Gaussian noise to the outputs of linear layers independently at every token position. Each token acts as a virtual population member, so one forward pass evaluates many perturbations in parallel; the reward-weighted noise is averaged into a gradient estimate. At a population of 16k, its estimates reach about 0.75 cosine similarity with backprop's gradients across most layer types.

Why does it matter?

At 10M tokens, Dust's test loss tracks backprop's closely (for example 5.095 vs 5.066 at 256 draws), and the authors report it is about 1,000 to 10,000 times more efficient than a transformer version of the EGGROLL evolution-strategies method. The authors state they do not try to make Dust compute-efficient enough to replace backprop today; it is a foundation for search-based credit assignment.

Who is it for?

ML optimization researchers

Try it

python dust.py --tokens 1M --population 256 --output runs/1m-p256

Sources · 3 outlets

Tags

  • dust
  • q-labs
  • zeroth-order-optimization
  • backpropagation
  • pretraining
  • transformers
  • evolution-strategies
  • optimization
  • research

← All releases · Learn AI