Robots Need Data, Collection Companies Proliferate: The Cost of Supply Chain Asymmetry
When 97 companies rush to collect "robot training data" using gimmicks like free cleaning services, VR games, and remote operation—it reveals a deeper truth: nobody actually knows which data has value, or how much it's worth.
8 min read
The Phenomenon: A Gold Rush and Gimmicks Everywhere
Between July 2025 and July 2026, nearly a hundred companies flooded China's embodied data collection sector, raising a combined 4.47 billion yuan. These companies spared no creativity: some set up "embodied data collection 5S shops" at China Mobile offices, letting customers collect data while doing chores; others packaged collection as VR games; some offered free home cleaning services; and still others built "cloud operation" platforms letting collectors remotely control robots.
These gimmicks keep multiplying, and while they appear creative on the surface, they actually reveal a deeper problem: nobody truly knows what data robots need, what quality of collected data is actually useful, or how much they're willing to pay for it.
Three Layers of Information Asymmetry
Layer One: Hidden Demand
AI companies like OpenAI, Google DeepMind, and Tesla need high-quality embodied data to train robots, but their definition of "high-quality" is often a black box. Collection firms only know the signal that "robots need data," but they don't know: - Out of 1 million hours of collected data, how many hours are actually usable for training? - Robots performing household tasks, industrial control, or medical care need completely different data characteristics, yet collection companies often apply one standardized process to all clients. - Will AI companies still buy data in the next three years? Or will they build their own collection teams?
Layer Two: Difficult Quality Verification
Data quality doesn't come with clear specifications like manufactured goods. The same 1 hour of data—"humans controlling robotic arms doing household chores"—can produce vastly different training value depending on: - Camera angle differences - Background complexity - Movement diversity - Annotation precision
But collection teams cannot know beforehand whether their data is Grade A or Grade C. They only find out after selling to AI companies and seeing the results in training.
Layer Three: Missing Price Discovery
Currently, there's no unified pricing for embodied data. Different companies, different use cases, and different quality levels produce wildly fluctuating prices. Some charge by hour, some by frame, some by task difficulty. This creates: - Collection companies unsure if their quotes are high or low - AI companies uncertain whether to build in-house teams or outsource collection - Investors unable to judge the true profit margins of the business
Economic Diagnosis: The Market for Lemons
Akerlof's classic paper shows: when buyers and sellers have significantly different levels of knowledge about product quality, low-quality goods drive out high-quality goods (bad money drives out good). Specifically:
1. Overcautious Selection: Uncertainty prompts AI companies to lower their price expectations for collected data, since they can't verify quality beforehand. 2. Oversupply: Seeing the funding wave, entrepreneurs rush in without understanding the true market size. 97 companies fighting over an ill-defined market inevitably creates overcapacity. 3. Excessive Innovation: Unable to compete through normal price mechanisms, firms resort to gimmicks—free cleaning, VR games, remote operation—all signals trying to tell buyers "we're different." But these gimmicks don't create data value; they just add costs. 4. Adverse Selection: The most patient collection companies, most willing to understand clients' real needs, get squeezed out for refusing to "burn money on gimmicks."
Why This Time Might Be Different
Embodied data collection isn't purely a lemon market, because three factors might break the asymmetry:
Factor 1: Technological Traceability If AI companies publicly disclosed "using this collection company's data improved model performance by X%," that creates a signal. But they haven't done so yet.
Factor 2: Long-Term Relationships Once AI companies and collection firms establish long-term partnerships, both gradually reduce uncertainty. But this requires time and depends heavily on early customers' patience for investment.
Factor 3: Standardization Efforts If industry associations or major purchasers (like Tesla) publish "embodied data collection standards," that effectively reduces information asymmetry. But no global standard exists yet.
Reality Scenarios
History shows three possible outcomes for lemon markets:
1. Market Contraction: Quality problems go unsolved, buyers exit gradually, and the industry shrinks to only large firms with extremely high reputation costs. 2. Information Intermediation: Independent quality verification agencies (testing bodies, rating firms) emerge, reducing information gaps between buyers and sellers. 3. Vertical Integration: Large AI companies (Tesla, Google) abandon outsourcing and build their own embodied data collection infrastructure, fully internalizing the data supply chain.
Tesla is already pursuing path 3—building its own data collection infrastructure for Optimus robot training. This may foreshadow that the 97 independent collection companies attracting 4.47 billion in financing will see their survival space dramatically compressed.
Not all data is worth collecting. Only when collection firms and AI companies reach consensus on "what is useful data"—and that consensus reflects reasonable long-term pricing—can embodied data businesses graduate from "gimmick competition" to "value competition."
Preparing your check…
Source: 36氪