A robotics data startup that most people hadn’t heard of in June is now in late-stage talks for a $1.2 billion valuation. According to TechCrunch AI, XDOF is negotiating a Series B led by 8VC, less than three months after it emerged from stealth. Sources with knowledge of the deal told TechCrunch that the company wasn’t even planning to raise again this soon.
This is significant because of the speed. XDOF closed a $70 million Series A in June with backing from Thrive Capital, Andreessen Horowitz, Lux, and Spark Capital. Now VCs are the ones knocking, not the founders. What flipped the script: annualized revenue approaching $50 million, per TechCrunch AI.
What XDOF actually does
XDOF collects real-world data for training general-purpose robots. Founded in 2024 by UC Berkeley researchers Philipp Wu (CEO) and Fred Shentu (CTO), the company builds the data pipelines, collection tools, and annotation systems that frontier AI labs and robotics firms can’t easily build on their own. Think of it as an outsourced data-supply chain for the robotics industry.
Investors are calling it the Scale AI or Mercor for physical robots, a nod to the data-labeling giants that helped power the current AI boom. The comparison points to why this matters.
Why robot data is the bottleneck
Large language models had a head start. They trained on the entire internet. Robots have no equivalent. There’s no giant, ready-made dataset of machines folding laundry or flattening boxes, which makes real-world data collection the critical bottleneck for building general-purpose machines.
XDOF is trying to close that gap two ways, according to TechCrunch AI:
- Remote teleoperation, where human operators steer robotic arms to generate training data
- Egocentric collection, where people wear body sensors to record everyday tasks like folding clothes and flattening boxes
The company plans to hire and train teams of data collectors worldwide for both roles. It’s also partnering with UC Berkeley’s AI Research lab to release what it believes is the largest collection of high-quality robot training data ever assembled, called ABC.
From a research problem to a company
The origin story is worth knowing. As a PhD student, Wu studied how robots learn from large datasets, and he kept hitting the same wall: not enough large-scale data to work with. So he and Shentu built GELLO, a low-cost teleoperation system that lets a human control a robotic arm remotely to generate training data. That project led to an influential robotics paper, and that paper became the foundation for XDOF.
The startup told TechCrunch it’s already working with 20 customers, including several frontier AI labs. That customer base, more than the valuation headline, is what explains the revenue and the investor rush.
What to watch next
A few caveats. TechCrunch AI reports it couldn’t confirm the total amount being raised or whether the $1.2 billion figure includes the new money. The terms aren’t final and could still shift. XDOF and 8VC didn’t respond to requests for comment.
Here’s what stands out to me. The money is chasing a thesis: whoever supplies the data for physical AI holds real leverage, the same way Scale AI did for the LLM era. XDOF isn’t alone in that bet. Rivals include Mecka AI, plus human-data platforms expanding beyond language models, such as Scale AI itself and Micro1. Expect that field to get more crowded and better funded fast.
For anyone building or investing in robotics, the signal is clear. Data collection is moving from an in-house afterthought to a specialized, venture-scale industry of its own. If XDOF closes this round at these terms, it sets a benchmark for how much that layer of the stack is worth. You can find the full details at the original source.