← All posts / Policy

Ten Days From Ask to Workstream: China's Data Regulator Moves to Write Embodied-AI Data Standards

China's National Data Administration will develop standards for embodied-AI training data, ten days after seven robotics firms asked for them — aiming at a 10-million-hour data gap.

Ten Days From Ask to Workstream: China's Data Regulator Moves to Write Embodied-AI Data Standards

Ten days. That is all it took for China’s National Data Administration (NDA) to turn a private industry request into an announced government workstream. On September 13, the agency said it will develop standards for embodied-AI training data and guide local authorities on the work — a direct answer to seven companies that had asked it for exactly that at a September 3 symposium. The move targets what has quietly become the single hardest constraint in humanoid robotics: not the models, but the hours of recorded reality needed to train them.

What the NDA announced

The statement, posted on the regulator’s website and first reported by Bloomberg and The Next Web, commits the agency to two things: writing data standards for embodied intelligence, and steering provincial and municipal authorities on how to support the work. Liu Liehong, the agency’s Party secretary and director, chaired a September 10 symposium attended by research institutes, technology companies, data firms and a humanoid training centre. The agency’s readout, published September 13, also pledged support for companies pouring money into data resources.

Notably, no draft standard, deadline, or public consultation window was disclosed. What was disclosed is the framing: “the industry is becoming more data-driven,” the agency said, and the seven companies it met on September 3 had “requested public data infrastructure for embodied AI and common data standards,” according to MLex’s earlier reporting on that meeting.

The speed is the signal. An industry ask on September 3 became a chaired symposium on September 10 and an announced standards workstream by September 13. In most regulatory ecosystems, that sequence takes quarters, not ten days.

The 10-million-hour problem

The gap these standards would sit on is enormous. By the China Academy of Information and Communications Technology’s (CAICT) estimate, embodied-AI foundation models need roughly 10 million hours of real-world training data. How much high-quality data of this kind exists worldwide today? Somewhere between 100,000 and 1 million hours — a shortfall of one to two orders of magnitude.

Language models had the internet. Embodied models have nothing equivalent. A robot learning to grasp objects, navigate cluttered warehouses, or manipulate fabric cannot download experience; it must record it, frame by frame, in physical space. That asymmetry is why the bottleneck in humanoid robotics has shifted from model architecture to data collection infrastructure — and why a data regulator, of all agencies, is now central to robotics policy.

China is already building the supply

The physical footprint is running ahead of the statistics. China operates more than 70 embodied-AI training grounds, spread across over half of its provincial-level regions, with 46 more planned or under construction, according to TechNode. About 86% of them target industrial manufacturing applications, and they cluster in three coastal belts: the Yangtze River Delta, the Beijing-Tianjin-Hebei region, and the Pearl River Delta.

These facilities exist to do one thing at scale: generate the hours. Each training ground is effectively a data factory — rows of robots attempting tasks under instrumented conditions, producing exactly the recorded interaction data that foundation models consume. The NDA’s standards push is best understood as the governance layer being fitted onto a collection apparatus that is already built and expanding.

The move also slots into an existing framework. In February, the Ministry of Industry and Information Technology’s technical committee released a six-part Humanoid Robot and Embodied Intelligence Standard System (2026 Edition). The NDA’s initiative adds data-specific rules alongside that MIIT framework — splitting the problem the way Beijing often does, with one agency owning the models-and-machines file and another owning the data file.

The contrast with Europe

Europe has written rules for this data without collecting much of it. The EU Data Act has applied since September 12, 2025, governing who may access the data a connected product generates — industrial machinery included — and prohibiting its use to build competing products. Design obligations bite on connected products placed on the EU market after September 12, 2026.

But the nearest European equivalent to a Chinese training ground is a single company: NEURA Robotics is building “gyms” — ten planned, five intended to be running by the end of this year, split between Europe, the United States and China. The Commission’s Data Union Strategy, published November 19, 2025 with data access for AI as its first priority, promised data labs that ten months on remain unestablished, with most of its actions not yet carried out.

The pattern is stark. One jurisdiction regulates access to machine-generated data it barely produces; the other produces the data at industrial scale and is now standardizing it. Neither approach alone wins — rules without data yield no robots, and data without rules yields no market trust — but the supply-side edge for embodied AI is drifting, facility by facility, toward China’s coastal clusters.

What to watch

The open questions are the binding ones. The September 13 statement disclosed no draft text and no consultation window, so how much force this acquires depends on whether the NDA moves from planning language to a dated instrument. Watch the agency’s next readout for a named deadline.

The second question is access: if foreign robotics firms want to train on data from the 70+ Chinese training grounds, will the coming standards permit it, or will they be scoped to domestic supply? An open standard attached to state-subsidized collection infrastructure would be one kind of geopolitical instrument; a closed one, another.

And a caution against triumphalism: only 23% of Chinese enterprises surveyed this year were satisfied with the robots currently on offer. A training ground collects hours, not paying customers. The data gap is closing on paper; the product gap is measured in deployment.

For now, the headline is procedural speed. Ten days from an industry request to a standards commitment is a pace few regulators anywhere can match — and in embodied AI, where the constraint is data, the jurisdiction that standardizes collection first writes the rules everyone else trains under.