OpenBMB Open-Sources Ultra-FineWeb-L1: 1.3T Tokens of 2025 Web Data Under Apache 2.0
OpenBMB releases Ultra-FineWeb-L1, a 1.3-trillion-token English web corpus built from six 2025 Common Crawl snapshots — the freshest open pretraining dataset to date, beating FineWeb by 0.6 points in ablation runs.
On August 20, 2026, OpenBMB — the open-source collective born from Tsinghua University’s ModelBest collaboration — quietly published one of the most consequential AI artifacts of the summer: Ultra-FineWeb-L1, a 1.3-trillion-token English web corpus carved out of six Common Crawl snapshots taken during 2025, released under an Apache 2.0 license on Hugging Face.
It sounds mundane. It is not. Pretraining data is the layer of the AI stack that fewest labs ever publish, and the open corpora that everyone trains on — FineWeb, FineWeb-Edu, DCLM, RedPajama — are all built from crawls that stopped in 2024 or earlier. Ultra-FineWeb-L1 is, to OpenBMB’s knowledge, the open web pretraining dataset covering the most recent Common Crawl snapshots in existence, reaching all the way to CC-MAIN-2025-51. If you are training an open base model in late 2026 and want your model’s world knowledge to include the second half of 2025, this is currently the only game in town.
What Ultra-FineWeb-L1 Actually Is
Ultra-FineWeb-L1 is the L1 filtered layer of OpenBMB’s UltraData framework, a tiered L0–L4 data management system the group introduced in February 2026. Within that hierarchy, L1 means “filtered”: basic cleaning, heuristic filtering, sensitive-field replacement, and deduplication — the heavy-but-not-selective processing that turns raw crawls into usable text. The tier above it, L2, is quality selection; L3 is refinement into synthetic Q&A and rewritten styles.
The first release contains 1T+ tokens across roughly 1.14 billion documents, organized by Common Crawl dump (CC-MAIN-2025-30 through CC-MAIN-2025-51) and shipped as Parquet files. Each document carries four fields: a UUID4 uid, the cleaned content, a JSON meta blob with URL, language score, WARC record ID and crawl date, and a dataset_index source identifier.
The pipeline behind it builds on Hugging Face’s FineWeb recipe but upgrades nearly every stage:
- Main-text extraction with trafilatura 2.0 — the newest major version of the extraction library, stripping navigation, comments, and boilerplate.
- fastText language filtering — retaining only high-confidence English documents.
- Heuristic filtering — FineWeb’s rules against repetition, boilerplate, and abnormal line structures.
- Sensitive-field replacement — emails, IPs, phone numbers, ID numbers, and credit-card numbers replaced with placeholders rather than simply dropped.
- MinHash deduplication — near-duplicate removal within each dump, following FineWeb’s per-dump strategy.
- Customized cleaning with agents — the most novel stage. OpenBMB used UltraData’s data-quality inspection tools and data-cleaning agents to hunt residual HTML, mojibake, invisible characters, corrupted content, and abnormal document lengths.
That last step is the tell: the boring parts of corpus construction are themselves being handed to language models now. The dataset that trains the next generation of models was, in part, cleaned by this generation.
The Numbers That Justify It
OpenBMB didn’t just dump the corpus — they ran the ablation. Following the FineWeb evaluation methodology, each data configuration was validated by training MiniCPM5-1B, a dense 1B model, for 20B tokens on identical settings: 32 GPUs, global batch size 512, Muon optimizer. Evaluation covers 12 benchmarks across six categories (general knowledge, reading comprehension, reasoning, NLU, math, table understanding) using 3-shot Cloze-format prompting.
On the shared snapshot CC-MAIN-2025-26 — the one snapshot both pipelines cover — the results at end of training:
| Data | Six-category Macro Avg | 12-task Micro Avg |
|---|---|---|
| FineWeb | 9.033% | 8.484% |
| Ultra-FineWeb-L1 | 9.668% | 9.180% |
| Ultra-FineWeb-from-FW | 9.954% | 9.414% |
| Ultra-FineWeb-from-L1 | 10.379% | 9.798% |
Read the table carefully, because it carries two findings. First, the upgraded L1 cleaning pipeline alone beats FineWeb by 0.635 points macro / 0.696 points micro on identical data — cleaner text, same crawl, better model. Second, the stack compounds: applying the Ultra-FineWeb classifier’s L2 quality selection on top of L1 cleaning (“from-L1”) yields the best results overall, outperforming selection applied on FineWeb (“from-FW”). Cleaning and selection are complementary, not interchangeable.
MiniCPM5-1B itself, released May 25, 2026, hit 1B-class open-source SOTA with Ultra-FineWeb as its core pretraining corpus — the ablation numbers are not a lab curiosity but the production recipe.
Why Fresher Crawl Data Matters More Than You’d Think
The obvious reaction to “1.3T tokens of web text” is that we already have plenty. Hugging Face’s FineWeb is 15T tokens. Nemotron, Qwen, and DCLM pipelines have published terabyte-scale corpora for two years. Size is not the story here — recency is.
A base model’s knowledge is frozen at its pretraining cutoff. Everything a model “knows” about software versions, current events, prices, people, and APIs comes from that moment. The open ecosystem has a quiet but serious problem: virtually every open pretraining corpus ends in 2024. Model builders who want 2025 knowledge have been forced to either license proprietary data, run their own crawl and cleaning pipeline (a multi-month engineering effort that FineWeb’s own logs show cost enormous compute), or distill from closed frontier models — with all the licensing and “model collapse” questions that implies.
Ultra-FineWeb-L1 changes that calculus. Six snapshots spanning mid-to-late 2025, cleaned with a validated pipeline, commercially usable under Apache 2.0. For the wave of open-weight models currently being trained — Meta’s Muse line, Alibaba’s Qwen 3.8 series, NVIDIA’s Nemotron push — a fresher open corpus directly attacks one of the last structural advantages of closed labs: proprietary, continuously-updated data pipelines.
There is a geopolitical footnote that’s hard to ignore. The two organizations releasing the most open pretraining data in 2026 are Chinese: OpenBMB (Tsinghua-affiliated) and Alibaba’s Qwen team. While Washington debates export controls on chips and “Pax Silica” bloc dynamics, the open data commons that everyone — including American labs — trains on is increasingly maintained from Beijing. Openness is turning out to be a strategy, not just an ideology.
The Fine Print
Two caveats deserve mention. First, the license has one unusual clause: while the code and pipeline artifacts are Apache 2.0, the README prohibits unauthorized unchanged redistribution — no mirroring or re-hosting the corpus wholesale without permission, a pragmatic hedge against带宽 costs and third-party repackaging that deviates from pure OSI norms. Derived use, training, and fine-tuning remain unrestricted, and users must still respect the rights of the original web sources, as with any crawl-derived corpus.
Second, “English, high-confidence only” is a scoping choice. The sibling releases cover other needs: Ultra-FineWeb (L2) provides ~1T English plus 120B Chinese tokens of classifier-selected data, and Ultra-FineWeb-L3 adds 400B+ English and 200B+ Chinese tokens of synthetic Q&A and multi-style rewrites — the latter still the largest open Chinese synthetic pretraining corpus.
Who Should Care
- Open-model trainers get a drop-in FineWeb replacement with fresher snapshots and a measured quality edge. The
from-L1result (10.379% macro) is the configuration to beat. - Data-engineering teams get a reference implementation of agent-assisted corpus cleaning — the pipeline is documented stage by stage, and the UltraData platform tooling is public.
- Researchers get something rarer: a controlled A/B experiment on cleaning-pipeline quality at the 20B-token scale, with identical training settings across four data configurations.
- Everyone else should note the date. The open-data commons just caught up to the end of 2025. The next generation of open base models will know things that happened after GPT-5 shipped — and that gap between open and closed, the knowledge cutoff itself, just narrowed by a year.
Ultra-FineWeb-L1 is available now at huggingface.co/datasets/openbmb/Ultra-FineWeb-L1, with the technical report on arXiv and the UltraData framework documented at ultradata.openbmb.cn.
Sources
- [1] https://huggingface.co/datasets/openbmb/Ultra-FineWeb-L1
- [2] https://huggingface.co/datasets/openbmb/Ultra-FineWeb
- [3] https://ultradata.openbmb.cn/
- [4] https://arxiv.org/abs/2505.05427
- [5] https://huggingface.co/openbmb/MiniCPM5-1B
- [6] https://buttondown.com/ai-tldr/archive/aitldr-daily-digest-august-20-2026/
- [7] https://huggingface.co/datasets/HuggingFaceFW/fineweb