This is the fourth part in a series on how AI is crossing out of software and into the physical world. This time, Jan Bosch examines the most expensive bet placed on top of that bottleneck: the human-shaped body.
The argument for the humanoid is the best argument in robotics, and I want to make it explicit and clear before I start dissecting it: The world is a brownfield. Every factory, warehouse, hospital and kitchen on the planet was designed around a particular machine: a bipedal primate about 1.7 meters tall with two five-fingered hands. Stairs, door handles, tool grips, workbench heights, the width of an aisle, the weight of a crate – all of it is a specification, and the specification is us. Build a robot to that specification and you have a universal adapter. No retrofit, no fixture, no re-layout. One machine, deployed anywhere, learning new tasks in software rather than steel.
That’s a genuinely powerful idea. It’s also the reason humanoids attract capital the way they do: The pitch is legible to anyone in about four seconds, which isn’t something you can say about many innovative ideas. We’re going to evaluate this idea from three perspectives. Not because I think the answer is obviously no (I don’t), but because I’d want clear answers before committing to a nine-figure initiative and I feel we’re not asking these questions consistently.
The first test is economic, and it’s about legs. A useful reference teardown of a full-size dexterous humanoid puts the bill of materials at roughly 55,000 dollars, with actuators alone accounting for about 56 percent of it. As the analysis puts it, a humanoid is “mostly a collection of expensive joints wrapped in structure.” Break the joints down by where they sit and the picture gets pointed. Locomotion hardware – thighs, calves, feet – comes to around 21,000 dollars, roughly 38 percent of the machine. The hands, which are the parts actually doing the work, come to 9,500 dollars, about 17 percent. The battery pack is 300 dollars.
Sit with that ratio for a moment. In the canonical humanoid, you’re spending more than twice as much on moving the robot between places as on the thing that manipulates the world when it arrives. If the task happens at one station, and in manufacturing, the overwhelming majority of tasks happen at one station, you’ve paid a 38 percent premium for a capability the task never uses, and taken on the balance, safety and power consequences of it as well.
The market already knows this, which is why the price ladder looks the way it does. The Unitree G1 sits at about 16,000 dollars and the R1 at 5,900 dollars, with the gap between them driven mostly by hand complexity and actuator count. Meanwhile, Chinese suppliers have undercut incumbent harmonic reducer manufacturers by 30–50 percent, which tells you where the cost curve is going and who’s going to own it.
The honest framing of the economic test is this: The humanoid isn’t competing against “retrofit the entire plant”; it’s competing against a purpose-built machine plus a modest fixture change. Run that comparison instead of the strawman and the universal adapter looks a lot less universal.
The second perspective is dexterity, and it extends the argument from the previous part in this series. Last time, I made the case that there’s no “internet of grasping.” Instead, manipulation data has to be produced by physical machines in real-time, one episode at a time. Rodney Brooks has been making a sharper version of that point, and I think he’s right in a way the field hasn’t absorbed. His argument isn’t that we lack enough data; it’s that we’re collecting the wrong kind.
“Collecting just visual data isn’t collecting the right data,” he writes. “There’s so much more going into human dexterity that visual data completely leaves out.” The human hand carries on the order of 17,000 low-threshold mechanoreceptors in the skin, firing as pressure is applied and released. Every one of them is a channel that the video of a person picking something up doesn’t record. Several of the best-funded humanoid programs are betting that watching humans on video is sufficient to learn what human hands do. Brooks’ position is that this is a category error rather than a scaling problem, and the burden of proof sits with the people spending the money.
I’d put it slightly more mildly than he does. We’ve seen apparently insurmountable data gaps close before. But a bet that requires a new sensing modality, a new data representation and a new collection apparatus to be invented is a different risk profile from a bet that requires more of what already exists, and it should be underwritten differently.
The third perspective is the one almost nobody brings up, and I think it’s the decisive one. Every industrial machine that has ever been certified to operate near people is safe for the same underlying reason: When something goes wrong, you cut the power and the machine stops. That single assumption, the fail-passive safe state, is the foundation under ISO 13849, under emergency stop categories, under the entire architecture of machinery safety. It’s why a robot can share a cell with a person.
A walking biped breaks it. A team of Siemens researchers put the problem about as plainly as it can be put in an August feasibility study: “Removing power from a walking biped produces an uncontrolled fall, which is itself a hazard.” The safe state of a humanoid isn’t de-energized; it’s an actively maintained balanced standstill that depends on the robot’s own real-time control policy continuing to function correctly. The thing you’d normally switch off to make the machine safe is the thing keeping it upright.
The consequences cascade in ways that will be familiar to anyone who has taken a system through certification. An aggressive protective stop can exceed the balance controller’s recovery margin and cause the fall it was meant to prevent, so the safety intervention becomes a hazard with no analogue in existing standards. Stopping time depends on where in its gait the machine happens to be. In the Siemens study, worst-case end-to-end response came out at around 1.1 seconds, dominated by the stopping dynamics rather than by detection or communication, and separation distances under ISO 13855 are computed from exactly that number. The authors, to their credit, explicitly decline to claim end-to-end PL e/SIL 3 certification. They frame the work as locating the gap precisely, and the gap is the robot’s own balancing policy.
Set that against the standards landscape. The 2025 overhaul of ISO 10218 folded ISO/TS 15066 into the main standard, added explicit functional safety requirements and, for the first time, admitted mobile platforms. Reporting on the revision indicates humanoids can now qualify as industrial robots on their manipulator characteristics while their mobility hazards sit outside the standard’s scope. Which means the single property that distinguishes a humanoid from a cobot on a pedestal is the property with no governing standard.
That, rather than any capability benchmark, is required before we can operationalize humanoid robots. Capability demos will keep improving and will keep telling us very little. The day a legged machine can be certified for mobility around people is the day the humanoid becomes an industrial product rather than a pilot.
For startups, the three viewpoints point somewhere specific. Whichever form factor wins, all of them need better actuators, cheaper harmonic reducers, tactile sensing that produces trainable data, and a functional safety architecture for actively balanced machines. Every one of those is a business that pays out regardless of how the humanoid question resolves, and at least one of them, safety for legged platforms, is currently a standards vacuum with well-capitalized customers walking toward it. Building the body is the bet that requires you to be right about the form; building the subsystem doesn’t.
If you’re building the body anyway, be honest with yourself about which claim you’re making. “A general-purpose machine will eventually be cheaper than many special-purpose ones” is a defensible long-horizon thesis. “A general-purpose machine is the right way to solve this customer’s problem in 2027” is usually not the same claim. As everywhere else in life, timing matters. Even if the claim is valid long-term, you can’t build a business that lives long enough to get to that future.
For incumbents, the temptation is to run a humanoid pilot for the wrong reason. Recall the number from the previous part in this series: BMW’s Figure 02 units accumulated roughly 1,250 operational hours across eleven months. That’s not a labor substitution; it’s a data collection program with a press release attached. Which is fine; Tesla is quite openly running Optimus the same way, provided the board knows that is what it has bought. The failure mode is a manufacturing executive who has been sold headcount reduction and receives an experiment.
The more consequential incumbent risk is procurement. If you deploy a machine whose distinguishing hazard falls outside the standard your safety case is built on, you’ve taken on a liability that your integrator’s certificate doesn’t cover. This will be discovered eventually, and it will be discovered by your legal function rather than your engineering function. If your supplier doesn’t provide the certification, you’ll have to do so instead.
For society, the humanoid does something no other machine in this series does: It looks like us. That triggers every intuition we have about replacement, and it does so in almost perfect anti-correlation with actual economic impact. The machines genuinely displacing work right now are shaped like conveyors, mobile shelving units and sorting arms, and nobody writes about them. The risk is that we end up regulating the silhouette rather than the capability: heavy scrutiny of an anthropomorphic machine doing 1,250 hours of pilot work and light scrutiny of the fleet quietly restructuring a warehouse.
There’s a real public-interest concern here, though, and it’s not the one that gets the coverage; it’s that machines with mass are being placed near people while the specific way they can fail, falling, has no governing standard. That’s a legitimate thing to be exercised about, and it’s considerably more actionable than the question of whether the robot has a face.
My own position, for what it’s worth: The general-purpose body is probably right eventually and wrong now, and the difficulty is that “eventually” is currently being financed at “now” valuations. The three perspectives are not arguments against humanoids; they’re the questions I’d want answered before treating one as a product rather than an experiment.
Louis Sullivan gave us the principle in Lippincott’s Magazine in 1896, and it has survived being quoted badly for well over a century: “Form ever follows function, and this is the law.” The humanoid asks us to run that backwards: to fix the form first, on the grounds that the world was built around it, and let the functions arrive later. That may still turn out to be the right call. It’s simply not the same kind of claim, and it should not be financed as though it were.


