NextFin News - Sunita Rathore sits atop a mountain of plastic waste on New Delhi's western fringe, an iPhone strapped to her forehead, recording every flick of her blade as she strips labels off discarded bags. She earns 150 rupees an hour for the footage — roughly $1.60 — and does not know it is feeding the artificial-intelligence models that may one day take her job sorting trash. Rathore's head-mounted camera is the front line of a quiet industrial revolution: thousands of Indian workers are being paid to film themselves doing physical work so that AI-powered robots can learn to do it instead.
The arrangement is the central paradox of the physical-AI boom. The same hands that stitch shoes, weld steel, slice mangoes, and segregate recyclables are generating the training data — called egocentric, or first-person, video — that robotics companies need to teach machines human dexterity. India, with its vast low-cost workforce and entrenched outsourcing infrastructure, has become the world's de facto data foundry for humanoid robots. But the work is a bridge with a destination written into it: every hour of video shortens the path to machines that no longer need the worker.
The Data Foundry: How Egocentric Video Became the Bottleneck for Humanoid Robots
Generative AI ate the internet — text, code, images, and video that already existed in digital form. Robotics cannot do that. There is no web-scale archive of what human hands do in the real world, and a robot that must fold laundry, load a dishwasher, or sort waste by grade needs exactly that: demonstrations of human motion in unstructured environments. A single-task policy for a narrow, controlled scenario may reach reasonable performance with 50 to 200 demonstrations, according to a 2026 survey of humanoid-robot data providers; a generalist household manipulation policy requires tens of thousands of diverse demonstrations across kitchens, objects, lighting conditions, and task variants. Companies including Figure AI and 1X Technologies are collecting data at the scale of hundreds of thousands of demonstrations for their generalist policies.
The economics of collection explain why the work migrated to India. Teleoperating a robot — having a human drive it remotely — is slow and expensive, tying one operator to one machine. Recording a human doing the task with a head-mounted camera is cheap, parallelizable, and produces footage that can train many policies. Research on imitation learning has shown that augmenting robot demonstrations with egocentric human data at roughly a 10-to-1 ratio improves performance while cutting the number of costly robot demonstrations required. The result is a global scramble for first-person video, and India is the cheapest place on earth to capture it at scale.
The pay tells the story. Rathore, a 43-year-old waste picker who makes about 20,000 rupees ($211) a month, earns an extra 150 rupees an hour for wearing the rig. In Chennai, Nagireddy Sriramyachandra, a 25-year-old housewife, earns 250 rupees ($2.60) an hour filming herself slicing mangoes.
"Who else will give you 250 rupees an hour just for doing housework?" asked Sriramyachandra from her kitchen.
Objectways, a data-collection firm with offices in Karur and Coimbatore in Tamil Nadu, pays 250 to 400 rupees per hour for egocentric capture — and still struggles to find enough workers.
The scale is already industrial. Ravi Shankar, president of Objectways, said the company is capturing 1,000 hours of data per day, against client demand for 200,000 to 300,000 hours.
"We started noticing this trend in mid 2025," Shankar said. "We are doing 1,000 hours of data per day, and the demand is for 200,000 to 300,000 hours of data."
Objectways began with GoPro cameras and Meta's smart glasses, then built its own recording app when the off-the-shelf hardware overheated during long shifts. The company works with Encord, which counts global robotics laboratories among its clients. Other players are moving in: Human Archive, a Y Combinator-backed startup founded by Berkeley and Stanford graduates, raised $8.2 million to build a distributed network for capturing egocentric, force-feedback, and teleoperation data, and has so far collected tens of thousands of hours with Asia — and within it, India — as its largest hub; Egolab.AI, founded in January 2026, bills itself as India's largest first-person point-of-view data aggregator and collects footage from garment-factory workers at suppliers including Pearl Global; Pronto, a Bengaluru home-services platform, has piloted outward-facing body cameras on its service workers.
For the workers, the income is real and welcome. Sriramyachandra said she is saving for the first time. Rathore is putting money aside for her four children's schooling. But the footage is captured under conditions that reveal the power imbalance at the heart of the trade: workers are told not to let faces appear on camera, they switch the device off when anyone calls out, and many do not know which companies will ultimately use the data or what it will build.
The Market: A Multi-Billion-Dollar Race Built on Human Labor
The data-collection boom sits on top of a humanoid-robotics market that is scaling from prototype to production. Estimates vary by methodology, but the direction is unanimous. One research house values the global humanoid-robot market at $3.93 billion in 2026, projecting $17.80 billion by 2031 — a 35.26% compound annual growth rate. Another puts 2026 at $10.9 billion and 2031 at $54.2 billion. Omdia forecasts more than 10,000 humanoid robots shipped annually by 2027, rising to 38,000 by 2030, an 83% CAGR from 2024. Tesla has targeted manufactured costs falling from $35,000 per unit in 2025 to $13,000 to $17,000 by 2030, with Optimus priced at $20,000 to $30,000.
Every one of those robots needs training data before it can do useful work, and that is where the annotation-and-labeling industry enters. The AI data-labeling market is estimated at $2.32 billion in 2026, up from $1.89 billion in 2025, and projected to reach $6.53 billion by 2031 at a 22.95% CAGR. Outsourcing already captured 54.85% of that market in 2025 and is growing at a 28.37% CAGR — faster than in-house operations — while video annotation is the fastest-growing data type, at a 31.18% CAGR through 2031. Manual workflows still accounted for 78.10% of labeling revenue in 2025, a reminder that automation has not automated the labelers yet.
The sector's most valuable private company illustrates the financial stakes. Scale AI generated an estimated $2 billion in revenue in 2025, up from $870 million in 2024, and carried a $29 billion valuation into 2026. Its business model — human-powered data labeling at scale, increasingly paired with automation tooling — is the template the egocentric-data vendors are following for physical AI. The difference is that robotics data is harder to fake: a model trained on bad footage produces a robot that drops objects, collides, or fails in the real world. Quality control requires humans watching humans, which keeps the labor loop closed even as the end product is designed to eliminate labor.
For India, the timing matters. The country's technology sector is navigating an AI-driven transition of its own. Nasscom expects the IT industry to add a net 135,000 jobs in fiscal 2026, taking total headcount to 5.95 million, with AI revenue from services firms of $10 billion to $12 billion — a small slice of an industry that has surpassed $300 billion. The egocentric-data trade offers a new export line, but one that monetizes the very capabilities — routine physical dexterity — that robotics is racing to automate.
The Mechanism: Why the Same Workers Build Their Own Replacement
The paradox is not an accident; it is the mechanism. Physical AI advances by imitation, and imitation requires demonstrations from the people currently doing the work. There is no shortcut around this. Simulation can generate variation, but sim-to-real transfer remains imperfect, and robots trained only in synthetic environments fail on the messiness of the real world — lighting changes, deformable objects, cluttered surfaces. The cheapest, highest-quality source of real-world demonstrations is the worker already performing the task for a wage. So the industry rents the worker's body for an hour, captures the motion, and uses it to train a policy that will eventually perform the task without the wage.
This creates a self-consuming loop. The more data a robotics company collects, the better its robots become; the better the robots become, the less human data they need for a given task, and the closer they come to displacing the data source. The workers are not being hired into a new occupation so much as converted, temporarily, into sensors. Their economic value in this chain is inversely related to the success of the product their labor creates.
The question investors should ask is whether this is a cyclical labor boom or a structural shift in the global division of labor. The answer is both, and separating them matters. The cyclical leg is straightforward: demand for egocentric data is surging because humanoid-robot programs are entering the data-hungry phase between prototype and deployment. That demand is real, growing at 30% or more annually, and it pays workers above their next-best alternative. It will persist as long as robot fleets remain small and policies remain data-starved.
The structural leg runs deeper and outlasts the cycle. Even after robots are deployed, they will need continuous data — for new tasks, new environments, and recovery from edge cases. Human-in-the-loop labeling is projected to grow at a 33.15% CAGR, and manual workflows still dominate at 78.10% share. But the structural endpoint is not mass unemployment overnight; it is a gradual compression of the wage share. As robots improve, the value of a given hour of human footage falls, because each hour trains a more capable machine. The workers who are paid to train the robots are, in effect, being paid to reduce the future market price of their own labor.
The Counter-Thesis: Why the Displacement May Take Longer Than the Headlines Suggest
The strongest case against the replacement narrative is that robotics has been five years away for thirty years. Dexterity, common sense, and safe operation in unstructured environments remain unsolved at consumer scale. Omdia's forecast of 38,000 humanoid units shipped annually by 2030 is substantial growth, but it is a rounding error against a global workforce of billions. A robot that can sort one recycling stream in one facility is not a general labor substitute. The capital cost, maintenance burden, and integration work mean deployment will concentrate in large warehouses and factories long before it reaches informal waste-picking colonies or home kitchens.
There is also a data-quality wall. Egocentric footage captured by low-cost head rigs is noisy: faces must be excluded, lighting is uncontrolled, and the camera moves with the worker's head rather than the hand doing the work — the very limb a robot needs to imitate. Cleaning, annotating, and structuring that footage into training-ready datasets is itself labor-intensive, which is why manual workflows still dominate the labeling market. If the data pipeline cannot scale cleanly, robot progress slows regardless of how many workers wear cameras.
Finally, the economics of substitution are not one-sided. A robot must be cheaper than the worker it replaces on a total-cost basis — purchase price, depreciation, energy, maintenance, supervision, and downtime — and reliable enough not to halt a production line. In low-wage settings, where a waste picker earns $211 a month, the robot has a very high bar to clear. Automation historically displaces tasks, not entire occupations, and it historically creates new work faster than it destroys old work in the aggregate. The International Monetary Fund estimated in 2024 that nearly 40% of global employment is exposed to AI, but exposure is not the same as displacement.
These counter-arguments are valid on the near-term horizon, and they should temper any claim of imminent mass job loss. But they do not overturn the direction of travel. The threshold does not need to be crossed everywhere at once; it needs to be crossed in enough high-value settings — warehouses, logistics hubs, assembly lines — to make the data flywheel self-funding. Once robot fleets reach the tens of thousands, the marginal cost of collecting additional data falls toward zero for the tasks the fleet already performs, and the substitution moves down the wage ladder.
What to Watch: The Signals That Will Settle the Debate
Three indicators will determine whether this is a temporary data gold rush or the beginning of a structural labor shift. First, the ratio of egocentric human data to robot-generated data in leading training runs. If human footage remains the dominant input through 2027 and 2028, demand for collection labor stays strong; if robot fleets begin generating most of their own training data through autonomous operation, the human-data boom peaks. Second, the cost per usable training hour. Objectways' 250 to 400 rupee hourly rate is a benchmark; if rates fall while demand holds, the work is becoming a commodity and the value is shifting to the model owners. Third, deployment density. Watch for announced fleet sizes crossing the 10,000-unit threshold at single operators — the point at which data collection begins to fund itself from operational use rather than venture capital.
For investors, the beneficiaries are asymmetric. The clearest winners are the data-infrastructure vendors and the model owners who capture the intellectual property, not the collection intermediaries competing on hourly labor rates. The exposed are the workers in routine physical roles — waste sorting, garment assembly, warehouse picking, home services — whose motions are the easiest to capture and the most valuable to automate. India's IT-BPM sector, employing nearly six million people, is already absorbing AI-driven pressure on entry-level white-collar work; the egocentric-data trade adds a parallel pressure on physical work, even as it temporarily employs some of the same population.
The base case is a multi-year boom in data-collection work layered on a steady decline in the long-run wage share of routine physical labor. The upside case is that robotics progress stalls on dexterity and common sense, leaving human data in demand for a decade or more and turning collection work into a durable, if low-paid, export industry. The downside case is faster-than-expected sim-to-real transfer and fleet-scale autonomous data generation, which would collapse the collection boom well before workers have transitioned to other work.
The falsifying signal for the displacement thesis is specific: if, by 2028, leading robotics laboratories report that less than half of their training data comes from human egocentric capture — because fleets are generating it autonomously — then the labor-replacement channel is weaker than it appears, and the data-collection boom is a bridge to nowhere for the workers. Conversely, if human footage remains the majority input while fleet sizes cross 10,000 units, the loop is self-reinforcing and the replacement timeline accelerates.
For Rathore and Sriramyachandra, the calculus is simpler. The camera pays today, and today is what matters when you are saving for a child's schooling or putting food on the table.
"I may get a robot myself in the future," Sriramyachandra said — a line that captures both the hope and the irony of the trade.
The workers training the robots are not ignorant of the risk; they are simply priced into it, one hour at a time. The physical-AI revolution will be built by the hands it is learning to replace, and the bill for that revolution will come due long after the cameras have come off.
Explore more exclusive insights at nextfin.ai.
