Blog

Quantifying Pretraining Variance: Data, Initialization, and Floating Point Arithmetic

We train paired 184M-parameter language models varying one factor at a time—initialization seed, data seed, floating point arithmetic order, and optimizer hyperparameters—and find that the variance induced by the ordering of floating point operations is within 20% of that from changing the data or initialization seed.