Three Months Out of Stealth, XDOF Is Already in Talks for a $1.2 Billion Series B
The robot-data infrastructure startup XDOF is in late-stage talks to raise a Series B at a $1.2 billion valuation led by 8VC, with annualized revenue already approaching $50 million and 20 customers including frontier AI labs.
Less than three months after emerging from stealth, XDOF — a startup that collects real-world teleoperation data for training general-purpose robots — is reportedly in late-stage talks to raise a Series B at a valuation of about $1.2 billion, with 8VC leading the round, according to a TechCrunch report published September 4, 2026.
The speed is the story. XDOF announced its $70 million Series A in June 2026, with participation from Thrive Capital, Andreessen Horowitz, Lux Capital, and Spark Capital. The company wasn’t planning to raise again so soon. But rapid growth — annualized revenue approaching $50 million — prompted venture capitalists to approach the company directly, several people with knowledge of the deal told TechCrunch. The terms are not final and could still change, and neither XDOF nor 8VC responded to requests for comment. TechCrunch could not learn the total capital being raised or whether the valuation includes the new funding.
From a Berkeley research problem to the ‘Scale AI of robotics’
XDOF was co-founded in 2024 by UC Berkeley researchers Philipp Wu (CEO) and Fred Shentu (CTO). As a PhD student, Wu studied how robots learn from large datasets. The biggest impediment to his research wasn’t compute or algorithms — it was the lack of “large-scale data to work with,” as he told TechCrunch in June. So he teamed up with Shentu on GELLO, a low-cost teleoperation system that lets a human operator remotely control a robotic arm to generate training data. The project led to an influential robotics paper, and that research became the foundation for XDOF.
The company’s pitch is infrastructure: XDOF builds the data pipelines, collection tools, and annotation systems that frontier AI labs and robotics companies can’t easily build themselves — in effect, an outsourced data-supply chain for the robotics industry. Investors now describe it as the “Scale AI or Mercor for physical robotics,” a reference to the data-labeling giants that fueled the LLM boom.
The name itself is a play on “degrees of freedom,” the robotics term for the number of independent motions a robot can perform.
Why robot data is the bottleneck
The comparison to Scale AI is more than a funding narrative. Large language models were trained on the entirety of the internet — trillions of tokens of text already sitting there, waiting to be scraped and cleaned. Physical robots have no equivalent. There is no internet of manipulations, no web-scale dataset of arms folding clothes, flattening boxes, or picking up objects in cluttered kitchens. Every trajectory has to be captured, one demonstration at a time.
That structural gap is what makes data collection a critical bottleneck for general-purpose robotics — and what makes a company whose core business is producing that data potentially as foundational to the physical-AI era as Scale AI was to the LLM era.
To capture it, XDOF combines two channels: remote robot teleoperation, where trained operators steer robots from afar, and egocentric collection, where humans wear body sensors to record everyday tasks like folding laundry and flattening boxes. The company plans to hire and train teams of data collectors worldwide, spanning both teleoperators and sensor-wearing egocentric operators.
ABC: the largest robot training dataset ever assembled
The company is also partnering with UC Berkeley’s AI Research lab to release what it believes is the largest collection of high-quality robot training data ever assembled, dubbed ABC. If the claim holds, a publicly released dataset of that scale would be a meaningful contribution to the field — the kind of shared substrate that could accelerate independent research the way open corpora did for NLP.
XDOF has previously said it is already working with 20 customers, including several frontier AI labs. Competitors include Mecka AI, another startup collecting real-world data for robot training, and human-data platforms expanding beyond LLMs, such as Scale AI itself and Micro1.
A unicorn valuation, with caveats
A $1.2 billion valuation on roughly $50 million in annualized revenue implies a revenue multiple around 24x — steep by traditional SaaS standards, but modest by the standards of 2026’s AI infrastructure frenzy, where startups with no revenue at all have commanded nine-figure valuations on team pedigree alone.
The bullish case: data supply is a durable moat in robotics. Foundation models for manipulation are racing ahead in capability, and every lab building them faces the same scarcity of real-world trajectories. A company that has industrialized the collection pipeline — warehouses, teleoperator workforces, annotation systems, quality control — occupies a chokepoint that gets more valuable, not less, as model capabilities improve.
The bear case: the round isn’t closed. “Late-stage talks” is a phrase that has dissolved before, valuations shift mid-negotiation, and TechCrunch’s sources could not confirm the raise size or whether $1.2 billion is pre-money or post-money. There’s also concentration risk — a handful of frontier lab customers could build collection operations in-house once volumes justify it, or standardize on a competitor. And the market is young enough that “the Scale AI of robotics” could still turn out to be several companies at once.
What it signals
Regardless of where the final terms land, the trajectory itself is the signal. A robotics data company going from stealth to near-unicorn valuation in a single quarter reflects where the industry believes the constraint now sits: not in robot hardware, which is maturing, and not in model architecture, which is increasingly shared, but in the physical interaction data that neither can substitute for.
The LLM boom created its own billion-dollar picks-and-shovels economy — labeling, evaluation, human feedback. The robotics boom is now visibly building the same shape of infrastructure around it, with human hands on real objects. XDOF’s Series B talks are the clearest evidence yet that investors think that layer, not the robots themselves, is where the durable value accrues first.
Sources
- [1] https://techcrunch.com/2026/09/04/xdof-just-three-months-out-of-stealth-is-in-talks-for-a-series-b-at-a-1-2b-valuation/
- [2] https://techcrunch.com/2026/06/17/collecting-robot-training-data-is-dirty-unglamorous-work-some-ai-labs-are-already-paying-xdof-to-do-it/
- [3] https://aiweekly.co/alerts/xdof-lands-70m-to-build-robot-training-data-pipelines
- [4] https://pulse2.com/xdof-raises-70-million-to-build-infrastructure-for-robot-foundation-models/