A robot at a company called Figure learned to fold a towel. It did not learn from a programmer writing rules about towels. It learned by watching 80 hours of video of people folding towels.
Just think about that number for a moment. 80 hours is two working weeks. In work terms, that is nothing. But it was enough to teach a machine a task that most people would struggle to explain in words.
Showing a machine what to do, instead of programming it, is how robots are being trained now. And it has created demand for a type of data that barely existed 5 years ago.
The industry calls it egocentric data. You probably know it as POV video.
Demand for POV video did not just grow in 2026. It exploded. Here is why robot companies want so much of this footage, and where all of it comes from.
What egocentric data means
Before we get to the trends, it helps to know what the words mean.
Egocentric data is video recorded from a person's own point of view. The camera sits on your head or your chest, so it sees what your eyes see. Your hands appear at the bottom of the frame, doing the work.
Researchers use a second term you will see often, which is exocentric. The difference is simple.
Ego means the self, so egocentric video is shot from the person's own viewpoint. Exo means outside, so exocentric video comes from a camera watching from across the room.
Both types are useful for training AI, but egocentric is the one robot companies are buying in bulk. There is a good reason for that.
Robots have no internet to learn from
To understand the demand, you need to understand what robots were missing up until now.
Large language models like ChatGPT learned from text and images that were already stored on the internet. Billions of pages, already written, already online, free to collect. The data was sitting there waiting.
Robots have nothing like that. There is no internet of physical actions. No website stores a record of how much force you use to pick up an egg, or how you angle your wrist to open a door. That information lives in human bodies, and until recently nobody had written it down.
Ken Goldberg, a robotics professor at UC Berkeley, put numbers on this gap in the journal Science Robotics. He calculated that the data used to train modern AI models adds up to roughly 100,000 years of experience:
Using commonly accepted metrics for converting word and image tokens into time, the amount of internet-scale data (texts and images) used to train contemporary VLMs [large vision-language models] is on the order of 100,000 years—it would take a human that long to read or view these data.
Ken Goldberg
The largest collection of robot training data ever assembled adds up to about 1 year.
He calls it the 100,000 year data gap. That single comparison explains most of what is happening in this market.
Karol Hausman, the chief executive of a robotics company called Physical Intelligence, described the same problem from the inside. His company cannot rely on free data from the internet, he said, because records of robot actions do not exist there. So the company has to create that data itself.
An independent research group called Epoch AI checked whether the problem was really about data or just about computing power. They found that robot models train on about 1% of the computing power used by leading AI models. Their conclusion was direct. Compute does not seem to be the blocker.
So the blocker is the data itself. And the fastest way to get it is to record real people.
4 reasons POV video beats a camera on a tripod
There are 4 reasons why the camera goes on your head rather than on a stand across the room.
1. Your head already points at what matters
The first and biggest reason is simple. You look at whatever you are working on, so a camera on your head looks at it too.
When you reach under a table to unhook something, you bend down and look at the hook first. The camera on your head goes down with you and records the hook up close.
Raunaq Bhirangi, a postdoctoral researcher at New York University, described this to IEEE Spectrum in August 2025. "The camera sort of moves with you," he said. "With this egocentric perspective, you get that information baked into your data for free."
Free is the important word there. Nobody has to mark up the video or decide which parts of it matter. The person recording already did that, just by looking.
2. A little footage goes a long way
Second, a robot does not need many hours to learn one action. Most tasks contain only a few seconds that really matter, so a short recording can hold everything the machine needs.
Bhirangi's team built a system called EgoZero and gave it 20 minutes of human video per task. The robot then handled 7 different tasks with a 70% success rate, without being trained on any robot footage first.
Training with no robot footage is the surprising part. A robot normally learns from recordings of other robots doing the job. This one learned from people alone, and 20 minutes per task was enough.
3. POV video scales where teleoperation cannot
Third, recording a person is far easier to scale than teleoperation, the method it replaces.
Teleoperation is where a trained operator steers a real robot by hand while a computer records every movement. It works well, but every hour of it needs a robot, a control rig, a trained operator and a prepared room. That puts a hard limit on how many hours the world can produce.
A 2026 study compared the two directly. Researchers took 5,000 hours of human POV video and 5,000 hours of teleoperation data, trained a separate model on each, then gave both models the same finishing training on real robots.
The models trained on human video scored 90% higher on tasks the robots had never seen before.
So the method that scales also produced the better robot. That is why companies stopped building recording labs and started recruiting people.
4. More footage keeps making robots better
Fourth, the gains do not stop. NVIDIA researchers kept adding hours of human video and measuring what happened, and performance kept improving at a steady rate. They found no point where extra footage stopped helping.
That is the finding behind everything else in this article. If more hours keep working, companies have a reason to keep buying them.
Video collections grew from thousands of hours to almost a million
In 2026 these video collections stopped looking like research projects and started looking like an industry:
Generalist AI trained its GEN-1 model on more than half a million hours of people handling real objects, recorded with low cost wearable devices. Five months earlier, at 270,000 hours, the company said that pile was growing by 10,000 hours a week.
NVIDIA built DreamDojo from 44,000 hours of human POV video, which it calls "the largest dataset to date for world model pretraining".
Peking University released HumanNet, a collection of almost 1 million hours of video of people.
Figure, the company with the towel folding robot, raised over $1 billion and said part of it would go on hiring people to record first person video. In September 2025 it also signed a partnership covering more than 100,000 homes, plus offices and warehouses.
China is doing this at national scale. Xinhua reported in March 2026 that 8 cities had opened data collection centres for this work. The retailer JD.com and the city of Suqian set a target of 10 million hours over 2 years.
Put those together and the pattern is clear. Every serious robotics company now treats recorded human activity as something it has to buy, and keep buying.
Wang Feili, a China analyst at UBS Securities, put it simply in January 2026. "Even if the 'baby' is born smart, without real-world datasets to feed it, it cannot grow."
Cooking, cleaning and repairs are what people film
None of this is dramatic footage. The value sits in ordinary activity, because ordinary activity is what robots are being built to handle.
Kitchen work is one of the most recorded categories. One of the first research datasets, EPIC-KITCHENS, was 100 hours filmed in 45 real kitchens, and cooking still appears in most collections today. Folding laundry, washing dishes, tidying and putting away shopping belong to the same group.
Repair and assembly work is the second large category. Anything done with hands and tools is useful, because grip, order of steps and small corrections are exactly what a robot has to learn.
Workplaces make up the third. Warehouses, shops, hotels and restaurants get recorded because the work repeats, and because those are the jobs robots are aimed at first.
Some programmes also ask people to say out loud what they are doing while they record. A spoken description gives the model a name for each action, so nobody has to add one afterwards.
How the footage reaches the AI labs
Very few of these companies record the video themselves. Collecting hundreds of thousands of hours means finding people in dozens of countries, giving them instructions, and checking every clip that comes back. An AI lab would rather spend its time training models.
So a supply industry grew up to serve them. Specialist platforms recruit the contributors, run the recording, review the footage, and sell the finished dataset to the labs.
We are part of that industry, and you should know it before you weigh up anything else in this article. Acquirox is our platform for supplying POV and egocentric datasets to the companies training these models.
It draws on the JumpTask network, which covers more than 18 million verified contributors across over 150 countries, and it records contributor consent before any work begins.
JumpTask offers POV data collection tasks, alongside other work that pays people to train AI. That means JumpTask earners are among the people producing this data.
We are not reporting on this market from the outside. We work at both ends of it.
Be one of the people recording it
JumpTask runs POV data collection tasks. Join free and see what is available where you are.
A phone and a head strap is enough equipment
You might picture expensive laboratory gear, and some of it is. Meta sells research glasses called Project Aria that carry 4 cameras, eye tracking and hand tracking in a frame weighing about 75 grams.
Most collection uses nothing like that. A university team published a full capture kit in 2026 costing about $151, made from 2 USB wrist cameras, some mounts, a head strap and a hub.
Reporters visiting collection sites have described people recording with phones strapped to their foreheads.
A phone and a head strap will do. That is the main reason this market grew as fast as it did.
A camera has no sense of touch
All that footage does not add up to a finished robot. Turning video into reliable physical skill is a separate problem, and nobody has solved it yet.
Rodney Brooks, who helped found iRobot and used to run the computer science and AI lab at MIT, argues that collecting visual data is not collecting the right data. Video captures what you see, but not what you feel.
It does not record how hard you gripped something, or how you corrected your hold when an object began to slip. That is why robots trained this way handle simple, open movements well and still struggle with delicate ones, like tying a knot or sealing a bag.
Even Goldberg, whose comparison opened this article, does not think more footage is the whole answer. His paper argues that traditional engineering has to do part of the work.
3 questions to ask before you record
Every hour in every dataset named above was recorded by a person doing something ordinary. The 44,000 hours, the half a million hours, the million hours: all of it came from people who agreed to switch a camera on.
That makes anyone who records a supplier in a large market, and it is worth knowing who you supply. Reporting through 2026 found programmes that never told people clearly what their footage would train, or who would end up holding it.
Ask 3 things before you start:
Who receives this footage, and who can they pass it on to?
What will it be used to train?
Can I stop, and can I have recordings deleted after I make them?
Then think about who else appears in your video. A camera on your head records anyone who enters the room, including family and children who agreed to nothing. A serious programme tells you where not to record and blurs the faces it captures by accident.
A programme that cannot answer all 3 questions in plain language is one to avoid.
What to take from this
Egocentric data is POV video, filmed from your own viewpoint.
Robots need it because there is no internet of physical actions. They have to be shown.
A head camera records what matters without being told. You look at your work, so nobody labels the footage later.
Small amounts work, and more keeps helping. 20 minutes per task reached a 70% success rate.
Demand is large and still rising. The biggest collections went from a few thousand hours to almost a million.
The activity and the equipment are both ordinary. Cooking, cleaning and repairs, filmed on a phone.
It is not solved yet. Video captures sight, not touch.
Ask before you record. Who receives the footage, what it trains, whether you can withdraw it, and where not to film.
Where this goes next
The collecting will not slow down soon. Companies in this market keep saying they need more hours, and researchers have not yet found a point where more stops helping.
The harder question is what all those hours produce. Recording is moving faster than the robots are, and the space between a full dataset and a useful machine is where the next few years will be decided.
Either way, something has already changed. Ordinary human activity is now raw material for AI, and the people who record it are the ones supplying it.
Ready to start recording?
Create an account, browse the tasks, and take the ones that suit you.
Rokas Brazinskas
Business development
Partnerships and Account Management professional with 6+ years of experience across B2B SaaS, eCommerce, and data infrastructure. Proven track record of driving MRR growth through strategic partnerships, API integrations, and high-value account ownership, managing portfolios of $100K–$150K+ MRR. Experienced in GTM execution, partner scaling, and bridging the gap between commercial and technical teams. Former people manager with international exposure.
Share:
IN THIS ARTICLE
What egocentric data means
Robots have no internet to learn from
4 reasons POV video beats a camera on a tripod
Video collections grew from thousands of hours to almost a million
Cooking, cleaning and repairs are what people film
How the footage reaches the AI labs
A phone and a head strap is enough equipment
A camera has no sense of touch
3 questions to ask before you record
What to take from this
Where this goes next
Make money online effortlessly
Get paid instantly for fun, easy tasks. No experience needed!