Appendix

A Architecture

The model is a decoder-only transformer of the modded-nanoGPT kind: eight pre-norm blocks of width 512 with eight heads, RMS normalization, rotary position embeddings with normalized queries and keys, a squared-ReLU MLP of width 2048, and no biases. Attention is causal with a sliding window of 1024 tokens in three of every four blocks and the full 2048 in the others. Every other block adds a value embedding to its values, a second token embedding scaled per head by a learned gate on the block input. The residual stream entering each block is a learned mix of the previous block’s output and the token embedding, and these $2L$ mixing scalars are the parameters trained by weight-space ES. The head is untied, with capped logits. Token and value embeddings are stored in bfloat16 and the forward runs in bfloat16. The base model has 37.7M parameters; the model-size study uses two, four and sixteen blocks of width 128, 256 and 1024, with the same head dimension.

B Cosine Ladder at the Remaining Checkpoints

Figure 6
Figure 6: The cosine ladder of Figure 5 at the checkpoints not shown there. Left, Dust at initialization. Middle and right, EGGROLL at initialization and at 100M tokens, each on its own axis.

C Test Losses by Token Budget and Population

Table 1 has the test losses behind Figure 2.

Table 1: Test loss at the validation selected checkpoint averaged over three seeds, across token budgets and populations. Population is draws rewarded with token losses for Dust and forward passes of the batch for EGGROLL (Section 3.1). Backprop is tuned on the same grid at every budget. The $\infty$ row is the limit of a power law fit in population.
Population100k1M10M20M
Dust647.2406.2905.3975.265
2567.1656.0675.2315.139
1k7.1515.9345.1114.960
4k7.1645.9115.0724.912
16k7.1705.9165.0494.802
$\infty$ (fit)7.1615.8925.0134.431
EGGROLL647.9617.4287.0566.987
2567.6057.2006.9066.811
1k7.3867.0466.6926.517
4k7.3026.8926.3526.200
16k7.2636.7086.0335.865
Backprop7.2165.9594.9894.633

D EGGROLL’s Extrapolated Population

We extrapolate EGGROLL’s ladder at 1M tokens and more to the population at which it would match Dust at 64 draws (Section 3.2). Table 2 gives that population as a multiple of 64 under three fits of EGGROLL’s cells in Table 1. The power law is the fit of Figure 2, $L(P) = A P^{-\alpha} + E$, which we also refit with each population left out in turn. The straight lines in log population are least-squares fits to the top three populations (1k, 4k and 16k) and the top two (4k and 16k). The three fits put the multiple between about 3,000 and 23,000.

Table 2: The population EGGROLL would need to match the test loss of Dust at 64 draws, as a multiple of 64, from fits of its five cells in Table 1. $\infty$ marks a refit whose floor lies above the loss of Dust.
Fit1M10M20M
Power law, all five populations23,000$\times$16,000$\times$8,100$\times$
Power law, one population left out11,000$\times$ to $\infty$7,200$\times$ to 54,000$\times$4,700$\times$ to 17,000$\times$
Straight line in log population, top three8,200$\times$3,700$\times$3,300$\times$
Straight line in log population, top two6,000$\times$4,100$\times$3,100$\times$

E Population Allocation

$K$ counts the draws rewarded with token losses. Table 3 lists everything an update jitters, with its draws at $K = 16{,}384$ and its noise scale, and Table 4 splits the three per-block streams over the blocks. Populations from 256 up scale every count by $K/16{,}384$. At 64 draws the recipe at 1M and 10M is $(8,7,4,3,1,1,1,1)$ writer draws, $(8,7,4,3,2,1,1,1)$ MLP draws, one attention-output draw per block, three embedding draws and 1, 5, 2, 3 and 2 draws for $q$, $v$, the gate, $k$ and the value embedding. The nominal stream ratio before rounding is $(\text{writers}+\text{MLP}):\text{attention outputs}:\text{embedding} = 7:0.75:0.25$ (at 100k tokens $6:1.5:0.5$), with a tilt toward the first blocks that comes from a coordinate descent over the allocation at 2k draws.

Table 3: What one update jitters, at $K = 16{,}384$. The first four rows are rewarded with token losses and add up to $K$ under SGD and to $11K/8$ in the Adam ladder. The SGD column is the allocation at 1M tokens, and Table 4 has the other budgets. The attention internals are credited through the attention output, and gates and value embeddings exist only on every other block. Each head draw jitters one slab of 256 logits. The attention outputs of the deeper half use 0.4 because their loss changes are weak and poorly resolved at 0.2. An update also runs one clean forward per GPU.
Draws
SiteOne draw rerunsSGDAdamNoise scale
Writer pairits block onward7,0404,0960.2
MLP hidden layerits block onward7,0404,0960.2
Attention outputits block onward1,6646,1440.2, deeper half 0.4
Token embeddingthe whole network6408,1920.2, Adam 0.4
Query, per blockits block’s attention1281280.05
Value, per blockits block’s attention1,2801,2800.05
Gate, per blockits block’s attention3843840.05
Key, per blockits block’s attention7687680.05
Value embedding, per blockits block’s attention5125120.05
Headone slab’s cross-entropy65,53649,1520.05
Residual scalarsthe whole network, twice880.03
Table 4: Draws per block at $K = 16{,}384$ for the writer pair (the MLP hidden layer gets the same, except block 2 at 100k, which gets 1,664) and for the attention output. The 16k cells at 10M and 20M have allocations of their own; every other cell scales its budget’s column by $K/16{,}384$. The token embedding gets 640 draws, and 1,152 at 100k. The 10M and 20M allocations at 16k take the same time per update as the one they replace.
Writer pair, MLP hidden layerAttention output
Block100k1M, 10M20M10M at 16k20M at 16k100kall others10M at 16k
120482432281621762304896384768
217922048243217922048768384640
38961024896896768512256512
4640640256512256384128256
5256384256384256128128256
6256256128256128128128256
7128128128128128128128256
8128128128128128128128256
Total61447040704062726016307216643200

Draws are evaluated $n$ at a time on one GPU, with $n = 1, 2, 8, 8, 16$ at $K = 64, 256, 1{,}024, 4{,}096, 16{,}384$, and split evenly across GPUs. Direct loss estimates and queries use $\gamma=0$. Keys, values, gates and value embeddings use $\gamma=0.99$ in the main comparison’s 10M/16k and 20M/16k runs, and $\gamma=0.98$ in all other runs. At 20M and 16k draws, $k$ and $v$ get 1,536 and 2,560 draws per block, so that cell has 14,336 draws rewarded with losses at the compute of the other 16k cells. For the other model sizes the tilt is sampled at the same relative depth and the attention-internal counts are scaled so that their totals match the 8-layer model.

F Tuning

Every cell of the main comparison, a token budget and a population for Dust and a token budget for backprop, is tuned the same way. Both methods tune momentum and learning rate only, starting from the grid of Table 5. Every other constant is fixed (Section 3.1, Appendix E), the learning rates of Table 6 included. Seed 42 runs the grid; the best point and every point within 0.02 of it get seeds 43 and 44, as does any point whose single seed beats the current three-seed best. A best point on the edge of the grid has the grid extended one step along its line until the best point is interior, and the two rate half-steps around it are then run under the same rules. A run whose final validation loss exceeds its best by more than 0.25, or that ever exceeds the uniform-prediction loss $\ln V$, or produces a NaN, is rejected together with its configuration. The selected configuration is the lowest mean validation loss over the three seeds, and we report its test loss at each seed’s best-validation checkpoint.

Table 5: The 12-point grid each cell starts from. The 64-draw cells at 10M and 20M need heavy momentum and small rates.
CellsMomentumLearning rate
100k tokens0.2, 0.3, 0.50.3, 0.4, 0.5, 0.6
1M, 10M and 20M tokens0.9, 0.95, 0.980.1, 0.15, 0.2, 0.25
10M tokens at 64 draws0.95, 0.98, 0.990.0125, 0.025, 0.05, 0.1
20M tokens at 64 draws0.95, 0.98, 0.990.00625, 0.0125, 0.025, 0.05
Table 6: Learning rates under SGD. Base is the tuned rate of the cell. The model-size study of Section 4.1 predates the protocol of Section 3.1, and its learning rate and momentum are selected per size and population.
Main comparisonModel-size study
ParametersDustBackpropDustBackprop
Key projectionsbasebase$4\times$ base$4\times$ base
Value projectionsbasebase$8\times$ base$8\times$ base
Token embedding100010000.3base
Value embeddings0.3base0.3base
Residual scalars0.03base0.03base
Everything elsebasebasebasebase

The overparameterization experiments were run before we improved the token-embedding learning rate. They use the earlier training settings in Table 6, which explains the difference from the main results.

Dust’s own constants, the noise scales, the credit decay and the draw allocation, are chosen with the cosine diagnostic of Section 2.3, confirmed in training, and held fixed across cells. At the 16k cells of 10M and 20M tokens they are tuned in training together with momentum and rate at fixed compute: each lever, the allocation across layer families, the per-block tilt, the noise scales, the credit decay and the head draws, is run one at a time from the (momentum, rate) winner, seeded in stages on wins paired by seed, and the winning levers are stacked; backprop receives the same wider (momentum, rate) grid at three seeds. The selected constants are those of Appendix E. Not every lever ran at 20M within the campaign, and the other cells keep the constants from the cosine tuning, so at those two budgets the top cell of the ladder has constants of its own, which the 10M fit is read with.

Under Adam, three learning-rate groups, matrices and the head, token and value embeddings, and residual mixing scalars, are tuned for Dust and backprop. Table 7 has the selected settings at 1M tokens.

Table 7: Selected settings under Adam at 1M tokens. Dust’s matrix rate is selected per population. All three methods use $\beta_2 = 0.95$, $\epsilon = 10^{-8}$ and no weight decay, and EGGROLL uses z-scored differences.
DustBackpropEGGROLL
Matrices and head, at 64 draws0.001750.0020.003
Matrices and head, from 256 to 4k0.00150.0020.003
Matrices and head, from 16k up0.0020.0020.003
Token and value embeddings$300\times$ matrices0.30.003
Residual scalars$100\times$ matrices0.20.003
$\beta_1$0.70.80.5

The model-size study of Section 4.1 changes two hyperparameters from Section 3 (Table 6). The token embedding stays at its default learning rate, 0.3 for Dust and the base rate for backprop, because at 1000 backprop shows no clear separation between model sizes, which makes overparameterization hard to study. We also keep the key and value learning rates of the earlier protocol and tune them for both Dust and backprop, so the comparison is fair to both.

G Cosine-fit parameters

Table 8 reports fits of $C(K) = c_{\max}/\sqrt{1+c/K}$ to measured cosine similarities at the final 100M token checkpoint. Each fit uses seven population budgets from 64 to 131072. At each budget, cosine similarity is computed separately for each weight matrix and then averaged across matrices of the same layer type. The fitted parameter $c$ uses the same population-budget units as $K$, rather than the individual draw count of a particular layer. RMSE is computed across the seven measured points used for fitting.

Table 8: Fitted parameters and errors for cosine similarity as a function of population size. Fits describe layer-type means.
Layer type$c_{\max}$$c$RMSE
100M-token checkpoint
MLP out0.8779316230.0081
MLP in0.8577484170.0065
Attn out0.8729216270.0093
Attn Q1.0000767360.0356
Attn K1.00006165950.0191
Attn V0.87761737800.0269
Value gate0.6657478630.0367
Value emb1.00004954500.0247
Token emb0.7318194980.0083
LM head1.00007590.0032

← Back to the post