Data Beat Architecture: The Six-Year Ablation That Rewrites Pretraining History
A granular 6-year ablation by Dwarkesh Patel and Jerry Han finds data improvements delivered 3.24x more compute-efficiency gains than model architecture from 2019 to 2025 — 12.0x for data versus 3.7x for models — with the two axes almost perfectly independent.
How much of the last six years of AI progress came from better data, and how much from better models? It sounds like a question someone should have answered long ago. Nobody had — until now. On September 8, 2026, Dwarkesh Patel (of the Dwarkesh Podcast) and Jerry Han published Pretraining Progress Is Mostly Coming From Data, a research write-up that decomposes six years of pretraining advances into their data and model components, and the headline number is blunt: from 2019 to 2025, 3.24x more compute-efficiency gains came from data improvements than from model improvements — a 12.0x compute multiplier for data versus 3.7x for models, measured at a fixed 10^19 FLOPs budget.
For an industry that allocates billions of dollars to architecture research, optimizer tuning, and kernel engineering — and treats data pipelines as unglamorous plumbing — this is an uncomfortable result. And it lands at a moment when the field is openly anxious about the “data wall,” the fear that the internet’s supply of fresh human text is running out.
What they actually did
The methodology is what makes this study credible where armchair arguments are not. The authors reconstructed, for each year from 2019 to 2025, two artifacts: a year-representative open model recipe (the publicly known algorithmic stack of that year — architecture, optimizer, initialization, learning-rate schedule, hyperparameters) and a year-representative public data corpus (the best open scrape plus curation techniques available at the time).
On the model side, the timeline runs from GPT-2 (2019) to OLMo-2 (2025), sweeping in RoPE positional encodings, RMSNorm, SwiGLU activations, QK-norm, parallel attention+MLP blocks, and the steady march of optimizer and initialization refinements along the way. On the data side, it runs from OpenWebText — roughly 9 billion tokens of upvoted-Reddit-linked web pages, which is essentially what GPT-2 trained on — to 2025’s UltraFineWeb-class corpora: whole-internet scrapes with classifier-driven filtering that learns to predict which documents will empirically improve model performance.
Then came the expensive part. The team trained the cross-product of these recipes and corpora — every model vintage against every data vintage — across five compute budgets from 10^17 to 10^19 FLOPs, with multiple independent seeds per point, using a shared tokenizer (GPT-2 BPE) and context length (T=2048) to keep comparisons clean.
There was one measurement wrinkle worth understanding: you cannot compare cross-entropy loss across runs trained on different datasets. So the authors evaluated end capabilities via OLMES (an aggregate of 10 relatively easy multiple-choice benchmarks) and derived “compute multipliers” — how much less compute a given recipe/corpus combination needs to hit a fixed capability bar, relative to the 2019 baseline. Noise from end-capability evaluation was handled with seed replications and parametric bootstrapping for error bars.
Three findings that matter
Finding 1: Data delivered 3.24x more efficiency gains than models. At the 10^19 FLOPs budget, the 2025 data corpus is worth a 12.0x compute multiplier on its own; the 2025 model recipe is worth 3.7x. Year over year, that decomposes into roughly 1.51x annual gains on the data axis versus 1.24x on the model axis, with a joint 1.57x yearly improvement when both are combined.
Finding 2: The two axes are almost perfectly independent. A linear model with purely additive model and data effects explains 88% of the variance in OLMES scores. Realizing the gains of OLMo-2’s architecture did not require UltraFineWeb specifically, and vice versa. That independence is quietly good news for the field: data work and architecture work are complementary levers, not competing ones, and neither is blocked on the other.
Finding 3: The caveat the authors are careful about — this is small-scale. Every run here is far below frontier scale. Model-side innovations that matter mostly at scale (MoE routing, sparse attention variants, stability tricks that keep hundreds of thousands of GPUs from diverging) don’t show up as compute multipliers at 10^19 FLOPs. The authors are explicit that inference-efficiency wins like GQA, tokenizer improvements, and long-context machinery fall outside what their metric can capture.
The sailboat and the container ship
The most valuable part of the write-up is the authors’ own refusal to over-claim. A naive reading says most 2019–2024 pretraining progress was “just data engineering” and model research barely mattered. They argue this is the wrong interpretation — and the analogy they reach for is nautical.
Small models are sailboats: they have little capacity, so you must be fanatically careful about what cargo you load. This is why data-quality gains dominate at their experimental scales. Frontier models are container ships: with enormous excess capacity and up to 100x overtraining relative to Chinchilla-optimal, they can absorb enormous volumes of lower-quality data and let stochastic gradient descent separate signal from noise. Aggressive filtering of a frontier corpus forces dozens of epochs over a smaller pile, which empirically performs worse than a larger, lower-average-quality dataset.
On this reading, model research’s historic contribution wasn’t efficiency — it was making scale usable at all. FlashAttention, MoEs, and normalization tricks didn’t primarily make training cheaper; they removed the constraints (exploding gradients, memory exhaustion, bandwidth walls) that made training across hundred-thousand-GPU clusters infeasible. Data engineering made each FLOP smarter; architecture made each cluster possible.
Why this matters now
The timing is what elevates this from interesting blog post to industry-relevant result. Three currents converge here.
First, the data wall. If data improvements have been the dominant driver of pretraining efficiency — and the corpuses studied are all curations of a finite Common Crawl stock, not expansions of it — then the main engine of pretraining progress depends on a resource that does not regenerate. “We’re not generating more internet,” as the authors put it. Whether synthetic data can genuinely expand the corpus without degrading models becomes, in their words, a crucial question — one they explicitly did not investigate and flag as future work.
Second, it quantifies a suspicion already in the air. Epoch AI’s trend tracking has estimated pretraining compute efficiency improving at roughly 3x per year, and Anson Ho and colleagues have conjectured that “most software progress might actually be due to data quality improvements.” This study is the first granular, bottom-up decomposition to attach hard numbers (12.0x vs 3.7x) to that conjecture — and notably, its measured joint 1.57x/year is below the 3x headline estimate, precisely because scale-dependent and inference-side gains don’t register at small scale. The gap between the two numbers is itself a research finding.
Third, it reframes the automation question. Ryan Greenblatt observed that most historical data-corpus improvements look like the kind of progress automated researchers could test empirically — run ablations on different data mixes, measure downstream performance, iterate. If that’s right, the dominant axis of pretraining progress is exactly the axis AI-driven R&D would accelerate first.
Honest limitations
The authors are unusually candid, and readers should hold both the findings and the caveats together: the scales are tiny relative to frontier training; OLMES is a battery of easy benchmarks that would reward different data engineering than, say, coding evaluations; the choice of one representative recipe and corpus per year is a judgment call, not an exhaustive survey; and their compute multipliers for some vintages (NeoX, the Pile) rest on extrapolation. The Pile underperforming OpenWebText on OLMES illustrates the point — its contribution was source diversity (PubMed, arXiv, GitHub, patents), which transfers minimally to English web-prose multiple choice.
None of which weakens the core takeaway. For six years, the field’s efficiency engine ran disproportionately on data. The next one may hinge on whether that engine’s fuel is renewable.