A quiet kind of extraction is underway in India, and as an expat it pains me to watch. It is not the old brain drain, where the talented left. This time the people stay and the value leaves.
Plumbers, electricians, factory hands, and domestic workers across remote and urban India are being fitted with specialized headgear that records their work, their movements, and their expertise. They are lured by an extra hundred rupees on the paycheck. What they are recording is the raw material for the robots that will eventually do their jobs. Automation of this kind of work is coming; that much is obvious and not the point. The point is where those robots get built and who owns their product. The data comes out of India, but the machines it trains will be designed, manufactured, and owned abroad. When the technology is ready, India becomes a subscriber to this technology, or will see Indian manufacturing growth stunted as manufacturing labor arbitrage vanishes and costs equalize globally.
An Indian factory worker is recording, for an extra hundred rupees, the skill that a foreign firm will package into a machine and sell back into India to make that worker redundant. The increase in wages is of immediate utility. The data it produces compounds enormously, into products and companies, and all of that accrues outside the country. On the surface it looks like a win: foreign investment arriving, cash reaching blue-collar workers who rarely see it, and even income for homemakers who get no pay for their daily chores. Underneath, it is a one-way pipeline. The skill and variety of Indian hands goes out cheap, and the value built on top of it stays out. This is the core issue, and it is why Indian leadership should treat it as a strategic question, not a labor story.
Why the data wants to be Indian
The reason this is happening in India specifically comes down to the shape of the data. Modern robotics AI needs the same kind of scaled corpus that let the LLM giants converge on superhuman performance, but robotics data has not scaled as much as text and media. The diversity of real-world physical tasks has never been captured in detail, and capturing it is not just a matter of quantity—it is about quality.
Quality is rarely collected directly. Approximations, augmentations, and conditioning are done extensively over a large corpus to arrive at reasonably good data. The real-world equivalent is asking 10 people for directions instead of one and extracting the parts that seem to match. India sits in a rare sweet spot: a deep pool of domain experts across countless niches, a low cost of labor, and enough general literacy to follow complex instructions.
There is a second advantage that is easy to miss. Most Western data is statistically near-uniform, clustered tightly around a standard, because high standardization means an electrician working a junction box in two counties of the same state produces nearly identical data. India is not like that. Every technician carries a body of knowledge that is partly their own, so the data comes back varied and full of the edge cases these models are starved for.
The harvest is already running
This is no longer hypothetical. Silicon Valley startups are already on the ground. Human Archive, founded by Berkeley and Stanford researchers and backed by 8.2 million dollars from Wing VC, Y Combinator, and angels at OpenAI, Nvidia, Google, and Meta, pays Indian gig workers a base rate of one dollar an hour to wear camera-equipped caps and sensor rigs while they clean, cook, and repair in people's homes. It reports more than 1,000 headsets already deployed. Competing outfits pay slightly more, roughly 250 to 400 rupees an hour, but the model is the same: cheap, large-scale capture of the exact egocentric data that physical-AI labs cannot get anywhere else. India's IT ministry has begun examining the consent mechanisms behind these programs, which tells you the stakes are already being recognized, if not yet acted on.
The demand is undeniable and scaling. Stellaris Venture Partners estimates that leading robotics labs will need between 100 million and 1 billion hours of egocentric data over the next two to three years, and one operator already collecting in India describes producing 1,000 hours a day against demand for 200,000 to 300,000 hours. A single training context, something as simple as picking up a glass and placing it on a shelf, can require 100,000 to a million hours on its own. The reason for such scale is due to the fact that what is simple for a child can be very challenging for a robot. The only way to achieve task parity or exceed human performance is to collect data at scale and at cost.
The foreign firms have local competition. Indian startups, Humyn AI, FPV Labs, Neo Cambrian, and others, are building the data pipelines themselves, and firms like Objectways are pivoting from LLM annotation into physical-AI collection. Whether the buyer sits in the Bay or in Bengaluru, the data flows outward to whichever lab can pay the most, and a Bay Area lab with hundreds of millions in funding will always outbid a domestic one. Humyn Labs alone runs a verified collection network spanning 18 countries across India, Latin America, Europe, and Southeast Asia. The pipeline is being built efficiently and at scale. The question is who ends up owning what comes out the other end.
This split is already visible in who gets hired where. Look at the open roles at a single well-funded robotics lab and the pattern is stark.
Figure 1. Open roles at Skild AI by location and function (June 2026), based on the public careers board. Roles open to both the US and Bangalore are counted as US, the likelier hiring location. What remains Bangalore-exclusive is mostly data-collection and robot-operations work, while the research, ML, and engineering roles that build the model sit in the US. Counts reflect the roles listed at time of access and are indicative of the pattern rather than exhaustive.
Two doors that look open but aren't
The usual pushback is that if not India, then Southeast Asia will supply the data. It is not that simple. Balancing many cheap annotators against a few expensive specialists is the everyday work of data scientists and engineers, and few places match India's combination of expertise, diversity, and cost. That blend is precisely what makes the data hard to source anywhere else.
But India cannot keep the value at home by outbidding them in an unregulated global market. Spending millions on specialized data is now a significant line item in every major robotics AI lab's budget, and there is no other reliable way to scale. A Bay Area startup with 10 million dollars or more in funding will outbid any Indian startup chasing the same data for itself.
We have been here before
None of this is unfamiliar territory. India has reached for the lever of controlling a strategic resource more than once, and the record is sobering. In 1978, the government drove IBM out of the country, enforcing a foreign-equity cap the company would not accept, with the hope that an Indian computer industry would grow in the space it left. That industry never arrived, but by sheer coincidence the software-services boom emerged, as the firms that filled IBM's void found Indian engineers cheap and capable and began routing work back home. India got something valuable even though it's not the thing it set out to build, and it settled the country into being the world's back office rather than an owner of the stack. India risks repeating this with data.
There are recent attempts which are even more instructive because they are closer to what I am proposing. In 2018, the Reserve Bank ordered all payment data to be stored inside India. The initiative was great, but execution speed was poor. Major companies like Visa and Mastercard sat on this for years, and it was not until 2021 that the digital infrastructure restrictions actually barred Mastercard from signing new customers. The 2019 Personal Data Protection Bill, which would have placed exactly this kind of localization control on Indian data, was withdrawn in 2022 after four years of deliberation. It collapsed on the one question, what data was actually covered? This left every business unable to estimate what compliance would cost and handed foreign entities an easy lever to push back. Taxing data would require very nuanced definitions to prevent ambiguity, but it must be generic enough to prevent abuse.
There is a counterexample of government infrastructure that worked. UPI worked. It did not work by taxing or banning Visa and Mastercard; it worked by building a public payment rail so good that the foreign networks became optional. It grew from 30 million users in 2017 to 400 million by 2024 and now moves money more than 170 billion times a year, and in doing so it gave India genuine sovereignty over its payment infrastructure. The lesson is not that India cannot defend a strategic asset. It is that India defends one by building the thing that captures its value, not by trying to wall it off.
ONDC, the open commerce network, is the more sobering version of the same lesson and the closer parallel. It has real scale, but its retail order growth did not survive subsidy cuts. It solved problems like fragmented ride-hailing and grew on its own, but it failed when it had to pry shoppers away from Amazon. The data backbone is closer to the first case than the second. Its buyers are a handful of labs and corporations already paying a dollar an hour for the supply, with no Indian alternative to defect to. The data rail India builds needs to stand on its own merits, not depend on subsidies.
Own the rails, don't wall them off
So here is what I propose, and it follows directly from that lesson. The instinct to treat this data as a controlled substance is right, but restriction alone is the part India has historically been bad at. The answer is to build, not just to wall off. India should stand up the equivalent of an NPCI for AI training data: a public entity that runs and governs a digital data exchange, beginning with the egocentric robotics data being harvested today and extending to every category that follows.
Crucially, the government should not collect the data itself. The IT ministry owns the rails; private startups and companies own the collection and the consumption on either side, exactly as banks and fintechs build the apps on top of UPI while NPCI runs the exchange underneath. That lets a domestic industry grow on both fronts at once. Providers feeding the network and subscribers building on it while the state stays out of operations it has no business running. And because the exchange runs through a public stack rather than a handful of private brokers, distribution can be made equitable by design. The same governance that sets consent and compensation floors also decides who gets access and on what terms, so the value need not flow only to whoever bids highest.
Clear definitions of the data and its cost of access lets the government control external access to ensure equitable pricing that does not fluctuate based on global capital allocation. The downstream data collectors and networks see limited downside because the IT ministry-governed exchange manages dynamic pricing transparently, preventing exploitation while giving domestic startups a structural advantage. No exploitation by bad actors, domestic startups get the best deals.
To the foreign labs, the IT ministry offers tax credits and priority/discounted data access if they set up in India and hire Indians into the critical research and engineering roles. The data cost then becomes a tax incentive instead of a penalty. The point is not to scare foreign investment off. It is to encourage seeding critical, technical, high-paying jobs in the country while also protecting against exploitative practices. This is the slow dividend that matters most, because every critical role filled here and every dataset owned at home is how a country stops supplying the labor for someone else's wave and starts owning the next one outright.
The precedent for this is now global. China has heavy regulations on what data can move out of the country. The UAE has set up protections that ensure there is no monopolization of regional Arabic linguistic and cultural data, the government actively backed massive internal data collection and processing. They built their own sovereign cloud infrastructure to ensure that local healthcare, public sector, and behavioral data never leave the country to train foreign-owned models. These examples prove that defending data sovereignty is now front and center for most commonwealths.
The next wave is the one that counts
India missed the AI wave, and that is okay. The transition from ML to AI right after COVID was something almost no one predicted; even people who worked on language models day to day could not foresee what the field would become in half a decade. But the good thing about waves is that there is always another one. What makes the Bay Area great is what I'll call wave-maxxing: the moment a wave breaks, a small but capable coalition declares it done and moves to anticipate the next one. AI for robotics is that next wave, and it has not had its GPT moment yet. The models exist, the training algorithms are mature, and the only missing piece is real-world data. India is sitting on that data right now.
This is the leverage, and it should not be squandered. The diversity and depth of Indian labor is a strategic asset on the scale of a natural resource, and right now it is being exported for a dollar an hour. It should instead be monetized deliberately, for the long term sovereignty of India's own technology stack. That does not mean another debate committee. It means a dedicated task force from the IT ministry paired with the country's technologists, watching for the next wave instead of the last one, and treating data not as an afterthought but as a national commodity. The wave is already here. India can build something that captures long term value, or the country can watch it pass the way the last one did.
References
- Ivan Mehta, "This startup is betting India's gig economy can train the world's robots," TechCrunch, May 26, 2026. techcrunch.com
- Swathi Moorthy, "Sometimes being egocentric is good. It can help train robots," The Economic Times, 2026. economictimes.indiatimes.com
- Skild AI careers board, open roles by location, accessed June 2026. job-boards.greenhouse.io
- "India's data localization efforts could do more harm than good," Atlantic Council, 2019. atlanticcouncil.org
- "The withdrawal of the proposed data protection law is a pragmatic move," Carnegie Endowment for International Peace, 2022. carnegieendowment.org
- "The tech revolution that wasn't," MIT News, 2026. news.mit.edu
- "India's UPI proves public digital model can surpass private networks," and IMF coverage of UPI as a global model, 2025–2026. orfonline.org
- Core42, UAE national sovereign cloud infrastructure. core42.ai