Dust: Pretraining Transformers Without Backpropagation
AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get school and study supplies delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Q Labs Research has introduced Dust, a zeroth-order method that trains transformer language models by perturbing activations rather than calculating backpropagation gradients. The October 2026 report says Dust performed competitively in its experiments, but its strongest results used larger populations and substantially more compute; independent validation and practical training costs remain unclear.

Q Labs Research has reported a method called Dust for pretraining transformer language models without backpropagation, the standard technique for calculating how model weights should change. In an October 2026 research report, the group says Dust’s zeroth-order approach was competitive with backpropagation in its tests, while also stressing that closer agreement with backprop required a larger population and substantially more compute.

Dust perturbs a model’s activations—the intermediate values produced as input passes through the network—rather than perturbing its weights. The method assigns a loss-based reward to perturbations and averages the reward-weighted changes to estimate an update. Q Labs says it perturbs activations independently at each token, treating tokens as members of a virtual population that can be evaluated together in a forward pass.

The report says Dust’s estimates align more closely with backpropagation as population size increases and remain well aligned across the scales tested, up to 1 billion tokens. It also reports that a 243-million-parameter model outperformed a model 120 times smaller at most population sizes. These are findings reported by the authors; the supplied material does not establish independent replication or results from production-scale language-model training.

Q Labs also compares Dust with EGGROLL, an evolution-strategy method that perturbs weights. Based on the report’s extrapolations, Dust is roughly 1,000 to 10,000 times more efficient than a transformer implementation of EGGROLL from 1 million tokens upward. That comparison is specific to the method and assumptions described by the authors; it is not a claim that Dust uses less compute than backpropagation.

At a glance
reportWhen: Published October 2026
The developmentQ Labs Research published an October 2026 report describing Dust, an activation-perturbation method for pretraining transformers without a backward pass.

A Different Route to Model Training

Backpropagation is central to modern neural-network training because it uses information about how a small change in each parameter affects the loss. Dust instead searches through activation perturbations and uses their observed effects to estimate updates. If such methods could train large models effectively, they could broaden the range of learning procedures researchers can test beyond those built around differentiability and backward passes.

The report does not show that Dust is a practical replacement for backprop. Its own account says larger populations bring estimates closer to backprop, which points to a trade-off between the method’s brute-force search and the compute required. The result is most relevant as an experimental demonstration that this approach can be tested on transformer pretraining—not evidence that current training systems should change.

Amazon

transformer language model training hardware

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

From Weight Search to Activations

Earlier evolution-strategy approaches perturb model weights and evaluate the resulting candidate networks. That can make large populations expensive because candidates must be represented and assessed. Q Labs positions Dust’s virtual population as a way to avoid separately materializing every candidate: perturbations are applied to activations during a forward pass, with different tokens serving as population members.

The report frames the work against the widespread reliance on backpropagation, which has shaped neural-network architectures, optimization methods and hardware. It argues that search-based methods might become more attractive as available compute grows. That is a motivation for the research, not a demonstrated forecast that brute-force methods will outperform gradient-based training across future systems.

“We present the first zeroth-order method that is competitive with backprop at pretraining transformer language models.”

— Q Labs Research, report summary

Amazon

GPU clusters for AI training

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Evidence and Compute Costs

The supplied report material does not give enough detail to establish how Dust’s training quality compares with backpropagation across full training runs, or whether the experiments have been independently reproduced. It also does not specify the total compute, wall-clock time, energy use or hardware required for the headline comparisons. Those details are needed to judge whether Dust could be practical beyond the reported tests.

The claimed efficiency advantage over EGGROLL is based on extrapolations, and should not be read as an advantage over backpropagation. It is also unclear from the available material how Dust performs across different data, model architectures and training objectives, or whether its results hold when scaling beyond the tested range.

Amazon

AI model training optimization tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Replication and Larger-Scale Tests

The report describes a research result rather than a scheduled product launch or a confirmed change to common training practice. The next evidence readers would need is independent replication, fuller compute accounting and direct comparisons with backpropagation under matched training conditions. Larger-scale tests would also clarify whether the reported population-efficiency pattern continues beyond the models and token counts covered by the authors’ experiments.

Amazon

neural network activation perturbation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Dust?

Dust is a zeroth-order training method from Q Labs Research. It estimates model updates by perturbing activations and measuring how those perturbations affect loss, rather than calculating backpropagation gradients.

Does Dust eliminate backpropagation?

The method is designed to train without a backward pass, but the report does not show that it has replaced backpropagation in routine or production-scale model training.

How does Dust use tokens as a population?

Q Labs says it independently perturbs activations at each token and treats each token as a virtual population member, evaluating the perturbations in parallel during a forward pass.

Did Dust use less compute than backpropagation?

The report says larger populations—and substantially more compute—help Dust’s estimates approximate backpropagation more closely. Its 1,000-to-10,000-times efficiency comparison is an extrapolation against EGGROLL, not against backpropagation.

Has the result been independently confirmed?

The supplied source is Q Labs Research’s report. It does not provide evidence of independent replication, so confirmation beyond the authors’ experiments remains unclear.

Source: hn

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

El Nino

Meteorologists confirm that El Niño conditions are developing, likely affecting weather worldwide. The event could influence droughts, storms, and agriculture.

Markets Are Competitive If And Only If P != NP

A recent theoretical breakthrough links market competitiveness to the P vs. NP problem, raising implications for economics and computer science.

Will The **High Temp In Philadelphia** Be >87° On Aug 3, 2026?

Market activity suggests uncertainty about whether Philadelphia’s high temperature will exceed 87°F on August 3, 2026, with recent trades indicating fluctuating predictions.

French Firefighters Face ‘Pyrocumulonimbus’ For First Time

French firefighters have reported experiencing a pyrocumulonimbus for the first time during wildfire suppression efforts, marking a significant development in wildfire behavior.