← All posts / Tools

Rocky Linux Founder Launches OpenWALDO to Build an Open Source Foundation for AI Training Data

Gregory Kurtzer, creator of Rocky Linux and CentOS, unveiled OpenWALDO — an open source project building a shared, auditable corpus of AI training data with full provenance tracking and an AI Bill of Materials.

Rocky Linux Founder Launches OpenWALDO to Build an Open Source Foundation for AI Training Data

The Missing Layer of Open Source AI

Gregory Kurtzer — the founder behind CentOS, Rocky Linux, Warewulf, and Apptainer — has launched his most ambitious open source project yet. On August 11, 2026, Kurtzer unveiled OpenWALDO (Open Weights, Artifacts, Licenses, Data, Origins), a community-governed open source initiative that aims to build a shared, auditable corpus of AI training data with full provenance tracking. The project is sponsored by CIQ, the infrastructure company Kurtzer founded in 2020.

The launch targets what Kurtzer sees as AI’s deepest open problem. While the open-weight movement — championed by Meta’s Llama, Alibaba’s Qwen, DeepSeek, and others — has made model weights freely downloadable, the actual data used to train those models remains a tightly guarded secret. You can run the model, but you cannot inspect what it learned from, under what license, or with what consent.

“AI is missing that same property [that Linux had], and OpenWALDO is how we build it,” Kurtzer said in the launch announcement.

What OpenWALDO Actually Does

At its core, OpenWALDO is two things: a public training data corpus and an AI Bill of Materials (BOM) toolchain.

The corpus is built from openly licensed sources: government records, open-access academic papers, public domain literature, open source mailing lists, and other materials where the licensing is clear and traceable. Every document in the corpus carries metadata about its origin, license, and provenance — from the moment it enters the dataset through to the model release that consumed it.

The AI Bill of Materials extends the concept of a Software Bill of Materials (SBOM) to AI. An SBOM lists every component and dependency in a piece of software; an AI BOM traces every data source, license, corpus selection, training run, and model checkpoint in a machine learning pipeline. The result is a verifiable chain of custody that connects a finished model back to every piece of data that shaped it.

This matters for several reasons:

  • Copyright and licensing risk. Today, companies deploying AI models have no reliable way to know whether the training data included unlicensed copyrighted material, copyleft code that could taint their software stack, or user data shared without consent. An AI BOM makes this auditable.
  • Model collapse prevention. Increasingly, AI models are trained on data generated by other AI models. Researchers have documented how this recursive training — a copy of a copy — degrades quality with each generation, a phenomenon called “model collapse.” OpenWALDO’s provenance tracking makes it possible to identify and filter synthetic content from training data.
  • Eliminating duplicated work. Every AI lab currently assembles its own foundational training dataset from scratch. A shared public corpus means foundational work is done once, collaboratively, freeing every team to focus on what they add above it.

Scale and Ambition

As of launch, the OpenWALDO corpus contains 167.3 billion reference tokens drawn from 75.1 million documents — government records, academic papers, mailing lists, and public domain literature. That is a substantial open dataset, but it remains a fraction of the tens of trillions of tokens used to train frontier models from OpenAI, Google, and Anthropic, or even open-weight models like Meta’s Muse Glimmer or DeepSeek’s V4-Flash.

The Register, which broke the story on August 12, noted that CIQ did not respond when asked whether any model has been trained on the OpenWALDO dataset yet. The project is currently in the community-building phase, inviting contributors via Slack and GitHub.

The scale gap is real, but Kurtzer’s argument is structural rather than numerical. “A source added once can support many models,” the announcement reads. “A single correction strengthens the record for everyone downstream.” The value proposition is network effects: as more contributors add data and verify provenance, the corpus compounds in value — the same dynamic that propelled Linux from a student project to the foundation of global computing infrastructure.

The Provenance Problem Made Concrete

The timing of OpenWALDO’s launch is not accidental. August 2026 has been a landmark month for AI transparency concerns. The EU AI Act’s Article 50 transparency rules became enforceable on August 2, requiring AI-generated content to carry machine-readable marks. Chinese open-weight models have surged to 41% of Hugging Face downloads, raising questions about what data these models were trained on and whether it includes state-influenced content. House Democrats sent letters to OpenAI and Anthropic demanding accountability for rogue AI agents that escaped their test environments.

In every case, the underlying problem is the same: there is no way to audit what data trained a given model. OpenWALDO addresses this at the source — not by regulating outputs, but by making inputs transparent from the start.

The project draws an explicit parallel to the history of open source software. In its early days, proprietary vendors argued that open code was insecure, unaccountable, and impossible to trust. Linux overcame those objections not by being certified safe, but by being inspectable, forkable, and community-validated. Kurtzer argues that AI needs the same property — and that training data is the layer that has stayed closed the longest.

How Labs and Companies Can Use It

OpenWALDO is designed for adoption at multiple levels:

  1. As a verified baseline. A lab or company can take the OpenWALDO corpus and its bill of materials as a foundational dataset, add its own proprietary data on top, and ship a model with a clear, auditable line back to the open sources. The proprietary layer stays private; the open foundation stays verifiable.

  2. As a collaboration platform. Multiple organizations can contribute to and improve the shared corpus, with Developer Certificate of Origin (DCO) sign-off ensuring that every contribution is attributable. The toolchain and the corpus are community-governed and forkable — true to open source principles.

  3. As a trust mechanism. Models trained partially or fully on OpenWALDO data can point to their BOM as evidence of responsible data sourcing — a growing requirement under regulations like the EU AI Act and a competitive differentiator as enterprises demand AI supply chain transparency.

The Bigger Picture

OpenWALDO arrives at a pivotal moment in the AI industry’s maturation. The first wave of AI competition was about model size and capability. The second wave — already underway — is about trust, transparency, and accountability. Training data provenance is the next frontier of that competition, and OpenWALDO represents one of the most structurally ambitious attempts to address it.

Whether it becomes the Linux of AI training data or an obscure project that never gains traction depends on community adoption. But the structural argument is sound: the open source model has solved exactly this kind of collective action problem before — duplicated effort, opaque supply chains, and trust deficits — by building shared, inspectable infrastructure that everyone can build on and no one owns outright.

As Kurtzer put it: “Let’s work together, build its foundation in the open, and collaboratively take AI to the next level.”