TL;DR
Get school and study supplies delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Q Labs Research has introduced Dust, a zeroth-order method that trains transformer language models by perturbing activations rather than calculating backpropagation gradients. The October 2026 report says Dust performed competitively in its experiments, but its strongest results used larger populations and substantially more compute; independent validation and practical training costs remain unclear.
Q Labs Research has reported a method called Dust for pretraining transformer language models without backpropagation, the standard technique for calculating how model weights should change. In an October 2026 research report, the group says Dust’s zeroth-order approach was competitive with backpropagation in its tests, while also stressing that closer agreement with backprop required a larger population and substantially more compute.
Dust perturbs a model’s activations—the intermediate values produced as input passes through the network—rather than perturbing its weights. The method assigns a loss-based reward to perturbations and averages the reward-weighted changes to estimate an update. Q Labs says it perturbs activations independently at each token, treating tokens as members of a virtual population that can be evaluated together in a forward pass.
The report says Dust’s estimates align more closely with backpropagation as population size increases and remain well aligned across the scales tested, up to 1 billion tokens. It also reports that a 243-million-parameter model outperformed a model 120 times smaller at most population sizes. These are findings reported by the authors; the supplied material does not establish independent replication or results from production-scale language-model training.
Q Labs also compares Dust with EGGROLL, an evolution-strategy method that perturbs weights. Based on the report’s extrapolations, Dust is roughly 1,000 to 10,000 times more efficient than a transformer implementation of EGGROLL from 1 million tokens upward. That comparison is specific to the method and assumptions described by the authors; it is not a claim that Dust uses less compute than backpropagation.
A Different Route to Model Training
Backpropagation is central to modern neural-network training because it uses information about how a small change in each parameter affects the loss. Dust instead searches through activation perturbations and uses their observed effects to estimate updates. If such methods could train large models effectively, they could broaden the range of learning procedures researchers can test beyond those built around differentiability and backward passes.
The report does not show that Dust is a practical replacement for backprop. Its own account says larger populations bring estimates closer to backprop, which points to a trade-off between the method’s brute-force search and the compute required. The result is most relevant as an experimental demonstration that this approach can be tested on transformer pretraining—not evidence that current training systems should change.
transformer language model training hardware
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
From Weight Search to Activations
Earlier evolution-strategy approaches perturb model weights and evaluate the resulting candidate networks. That can make large populations expensive because candidates must be represented and assessed. Q Labs positions Dust’s virtual population as a way to avoid separately materializing every candidate: perturbations are applied to activations during a forward pass, with different tokens serving as population members.
The report frames the work against the widespread reliance on backpropagation, which has shaped neural-network architectures, optimization methods and hardware. It argues that search-based methods might become more attractive as available compute grows. That is a motivation for the research, not a demonstrated forecast that brute-force methods will outperform gradient-based training across future systems.
“We present the first zeroth-order method that is competitive with backprop at pretraining transformer language models.”
— Q Labs Research, report summary
As an affiliate, we earn on qualifying purchases.
Evidence and Compute Costs
The supplied report material does not give enough detail to establish how Dust’s training quality compares with backpropagation across full training runs, or whether the experiments have been independently reproduced. It also does not specify the total compute, wall-clock time, energy use or hardware required for the headline comparisons. Those details are needed to judge whether Dust could be practical beyond the reported tests.
The claimed efficiency advantage over EGGROLL is based on extrapolations, and should not be read as an advantage over backpropagation. It is also unclear from the available material how Dust performs across different data, model architectures and training objectives, or whether its results hold when scaling beyond the tested range.
AI model training optimization tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Replication and Larger-Scale Tests
The report describes a research result rather than a scheduled product launch or a confirmed change to common training practice. The next evidence readers would need is independent replication, fuller compute accounting and direct comparisons with backpropagation under matched training conditions. Larger-scale tests would also clarify whether the reported population-efficiency pattern continues beyond the models and token counts covered by the authors’ experiments.
neural network activation perturbation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is Dust?
Dust is a zeroth-order training method from Q Labs Research. It estimates model updates by perturbing activations and measuring how those perturbations affect loss, rather than calculating backpropagation gradients.
Does Dust eliminate backpropagation?
The method is designed to train without a backward pass, but the report does not show that it has replaced backpropagation in routine or production-scale model training.
How does Dust use tokens as a population?
Q Labs says it independently perturbs activations at each token and treats each token as a virtual population member, evaluating the perturbations in parallel during a forward pass.
Did Dust use less compute than backpropagation?
The report says larger populations—and substantially more compute—help Dust’s estimates approximate backpropagation more closely. Its 1,000-to-10,000-times efficiency comparison is an extrapolation against EGGROLL, not against backpropagation.
Has the result been independently confirmed?
The supplied source is Q Labs Research’s report. It does not provide evidence of independent replication, so confirmation beyond the authors’ experiments remains unclear.
Source: hn
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
