This is the third part in a series on how AI is crossing out of software and into the physical world. This time, Jan Bosch reaches the hinge: What happens when a model has to act?
The comforting story this time is a story about transfer. Vision is solved. Language is solved. Robotics, with AI going embodied, is simply an integration problem: Take a model that already understands the world, connect it to an arm and a camera, and the competence flows downhill into the machine. The hard part was the intelligence, and we have that now.
I spend a lot of my time with companies that build things with mass and momentum and have seen time and again how complex it is to combine mechanics, electronics and software into an integrated system. And if you’ve worked on the same type of systems, I think you’ll find the above story as hard to believe as I do.
Let me start with vision-language-action models (VLAs). VLAs are a substantial innovation, they work and they’re moving fast. Google Deepmind’s Gemini Robotics On-Device 2, published in July, takes a text instruction, images from the robot’s own viewpoint and proprioceptive data, and outputs actions. Crucially, it runs on the robot. No round trip, no dependency on a network. Robotics is the domain where the edge stops being a cost optimization and becomes a physical requirement, especially when there are safety concerns.
So, the capability is arriving. The problem is the training data. A language model is trained on the accumulated written output of humanity, which was produced for other reasons and happens to be lying around. There’s no equivalent corpus for manipulation. Every training example of a robot picking up an unfamiliar object in an unfamiliar pose has to be produced by an actual machine, in actual time, usually with a human teleoperating it. The data doesn’t exist until someone pays for a robot and an operator to generate it, one episode at a time, at roughly the speed of physical reality.
This isn’t a temporary shortage that scale will resolve; it’s a structural difference in how the two kinds of models get their competence, and it shows up in the results. A recent analysis of 1,228 vision-language-action papers published between February 2023 and June 2026 puts numbers on something the field has been quietly uncomfortable about. On the Libero benchmark, success rates that sit around 95 percent collapse to below 30 percent under modest perturbations of the scene. On another benchmark, performance went from over 90 percent to exactly zero. Not degraded. Zero.
Robots are great at repeating the exact same pattern, but even small changes tend to degrade behavior below acceptable norms. It’s not about a model that has learned to manipulate objects and is having a bad day; that’s a model that has memorized a set of trajectories and is being asked to do something fractionally outside them. And the honest players say so themselves: The Gemini Robotics On-Device 2 model card states plainly that it has “limited ability” to generalize to out-of-distribution tasks.
I want to be careful here, because this is easy to misread as robot pessimism, and it’s not. Robots work. They work extraordinarily well. The International Federation of Robotics counted 542,000 industrial robots installed in 2024 and an operational stock of 4.66 million machines worldwide, up 9 percent in a year. Those machines weld, place and palletize with a reliability that no language model comes close to. They achieve it by not generalizing at all. They’re programmed against a fixture, a part and a cell, and everything outside that envelope is engineered away rather than learned. This is traditional automation, not embodied intelligence.
The question isn’t whether robots work; it’s whether general-purpose robots do, and there, the evidence is thinner than the funding implies. The most instructive number I’ve seen this year is from BMW’s deployment of Figure 02 units, which accumulated approximately 1,250 operational hours over eleven months. Take a moment with that. Eleven months, and the machines were working under four hours a day across the calendar. For an industrial asset, that’s not a deployment; it’s an extended and expensive experiment. Meanwhile, venture capital put 40.7 billion dollars into robotics in 2025, and Bank of America projects on the order of 90,000 humanoid shipments in 2026. The gap between the capital and the operational hours is the whole story.
Of course, the first versions of ChatGPT were quite limited in usefulness and the progress since then has been phenomenal. With all the VC investment into robotics, we’ll see progress in this area as well, but LLMs have a corpus of training data that robotics can only dream of.
For a startup, this changes what you’re actually building. If the model is downloadable and the hardware is increasingly purchasable, then where’s your differentiation? To me, the answer is the apparatus that generates data in your specific slice of physical reality, and the tighter that slice, the better your odds.
Although the promise is general-purpose robots, the right answer is, in my view, exactly the opposite. Generality is a claim about a distribution and the only way to close a distribution is to bound it. A startup that picks one task, in one class of environment, with one gripper, can plausibly collect enough trajectories to cover the tail. And the tail is where all the value and all the liability sit. A startup that promises a robot that does everything has committed to a data collection problem it can’t finance. The correct question at the seed stage isn’t “How capable is your model?” but “How many hours of real-world interaction do you need before the failure rate is acceptable, and who’s paying for them?”
Think about autonomous cars, which could be viewed as a narrow, specialized type of robot. After well over a decade of promises, we’re still waiting for general availability. The solution was careful training in one city or even part of a city, as Waymo and others are doing. Or in logistics, autonomous trucks that only drive one stretch between a warehouse and a factory. The answer is to go as narrow as you can and still have a business case.
There’s a trap here worth naming explicitly, because the Libero numbers make it concrete. A demo is a sample from the training distribution. It tells you almost nothing about performance one step outside it, and the collapse from 90 percent to zero isn’t a gradual slope you can extrapolate along. Any investor or executive evaluating a robotics company should be asking to see the perturbation results, not the highlight reel. If a team can’t show you what happens when the lighting changes and the object is rotated forty degrees, they haven’t measured the thing that matters.
For large incumbents, the calculus inverts again, and in this case, it inverts in an unusually interesting way. Those 4.66 million installed machines represent the largest corpus of physical interaction data in existence. And almost none of it is being captured in a form that could train anything. Every one of those robots has been executing, logging and correcting for years. Whoever works out how to instrument an installed base and turn decades of industrial motion into training data has a genuine asset that no amount of venture funding can replicate, because the robots are already bolted to the floor.
The risk on the other side is disintermediation at the intelligence layer. If the model becomes the thing that determines what a machine can do, then a company that manufactures excellent arms and buys its intelligence from someone else has become a supplier of commodity actuation. For both incumbents and startups, it’s critically important to understand where the value sits going forward. As decades of digitalization have taught us, the value shifts from atoms to bits, no matter whether it’s software, data or AI.
There’s also a geographic fact here that echoes the first part of this series. Asia took 74 percent of new robot installations in 2024, China alone accounted for 295,000 units and Chinese domestic manufacturers now hold 57 percent of their home market, up from 28 percent a decade ago. The compute substrate is concentrated in a few places, and so is the embodiment layer – just not the same few places. Still, if the training data is the key restriction, I think it’s obvious where the most data from new robot installations can be collected and, once again, it’s not in Europe.
For society, the reflex is to go straight to employment, and that conversation is worth having, but it’s not the one I find most pressing. The more immediate issue is a mismatch in what we’re prepared to tolerate.
We’ve collectively decided that a language model being wrong some of the time is acceptable, because the cost of a wrong answer is that someone reads it and moves on. That tolerance doesn’t survive contact with a machine that has mass. A humanoid working near a person can’t be right 95 percent of the time; the production bar has to be well north of 99.9 percent. The difference between those two numbers isn’t incremental engineering but a different discipline entirely. Physical AI takes a technology whose defining characteristic is probabilistic behavior and puts it in a setting that has spent a century engineering probabilistic behavior out.
Even in contexts where no humans can be harmed because of probabilistic behavior, the financial consequences of robots damaging themselves, other robots, products or surrounding infrastructure simply destroy the business case. Systems that interact with the physical world need to be engineered with completely different reliability, robustness and safety requirements.
This is where the thread I’ve been pulling on since the summer becomes unavoidable. When the realization is a learned policy rather than written code, what you version and defend is the contract the machine must honor plus the running evidence that it still does. For a physical system, that contract has to specify not only what the machine will do but what it will never do, and the evidence has to be continuous rather than collected once at commissioning. We don’t yet have good engineering practice for this. We have benchmarks that collapse to zero under a rotation, which is roughly where software testing was before anyone thought to write down what a regression was.
The optimistic reading is that manipulation is a data problem, and data problems are the kind our industry knows how to attack once it stops pretending they’re model problems. The pessimistic reading is that the tail of physical reality is longer than anyone’s capital. My own view sits closer to the first than the second, but with the timeline stretched: Narrow, bounded, well-instrumented deployments will compound quietly for years before anything deserving the word “general” shows up on a shop floor.
Which is a modern restatement of something a roboticist worked out before most of this industry existed. Hans Moravec, in “Mind children” in 1988: “It’s comparatively easy to make computers exhibit adult-level performance on intelligence tests or playing checkers, and difficult or impossible to give them the skills of a one-year-old when it comes to perception and mobility.” We built the adult; we’re still working on the one-year-old.


