Back to blog

The 95-to-60 Cliff: Reading the Humanoid-Robot Boom Without Believing the Keynote

9 min read
Illustration: a generic humanoid robot in a warehouse aisle, scan-lines from its head parsing a shelf, one box highlighted in amber
The model can see the box, name the box, and plan to lift the box. Doing it reliably, ten thousand times, is the whole remaining problem. Illustration: AI-generated.

2026 is the year embodied AI stopped being a demo reel and started being a logistics line item — and also the year the gap between the press release and the loading dock became impossible to ignore. Both things are true at once, and holding them together is the only honest way to read the humanoid-robot boom. Let me try.

The technical unlock is real and worth naming precisely: Vision-Language-Action (VLA) models. These take camera input and a natural-language instruction and emit motor actions directly, learning generalizable policies the way LLMs learned generalizable text. ICLR 2026 had 164 VLA submissions; the architecture has clearly become the dominant paradigm, following the exact trajectory language models blazed a few years earlier. That is not hype — it is a genuine methodological shift, and it is why the field feels like it inflected this year rather than merely improving.

What is actually deployed — and how we know

Here is where discipline matters, because the deployment numbers are a minefield of unaudited vendor theater. Read them by source quality:

  • Agility Robotics reports its Digit robot moved 100,000+ totes in a commercial deployment at GXO. This is the kind of claim I weight heavily: a specific, operational metric, tied to a named customer, in a real facility.
  • Figure AI is reported to have surpassed 10,000 deployments across partner warehouses. Softer than a throughput number, but plausibly grounded in real installations.
  • Tesla Optimus is reported to have passed 50,000 cumulative units — a figure Tesla itself has never published or audited. On the Q4 2025 call, Musk conceded Optimus was "not in usage in our factories in a material way." As of the reporting: zero external customers, zero verified productive factory deployments.
  • Unitree ships more humanoids than any Western competitor at roughly a tenth of the price — and still saw Q1 2026 profit fall by half, which tells you the unit economics are nowhere near settled.

Notice the pattern: the most impressive number (Tesla’s 50,000) has the weakest provenance, and the most credible claim (Agility’s totes-moved) is the least glamorous. That inversion is the entire information hazard of this sector. A cumulative unit count is a manufacturing brag; totes moved in a customer’s live operation is a productivity fact. Learn to want the second kind and discount the first, or you will build a strategy on a keynote.

The 95-to-60 cliff is the real story

The single most important number in embodied AI right now is not a deployment count. It is this: policies that work 95% of the time in the lab drop to about 60% in live operations. That gap — the sim-to-real, demo-to-deployment cliff — is where every honest robotics timeline lives or dies.

Why it matters so much: a 95%-reliable warehouse robot sounds nearly done. A 60%-reliable one is worse than useless, because a robot that fails four times out of ten needs a human babysitter, and now you are paying for a robot and a supervisor to do one person’s job. The economics of automation are brutally nonlinear in reliability — value is close to zero until you cross some high threshold (often cited near 99.x% for lights-out operation), and then it snaps to enormous. The demos are real. The demos are also specifically the 95% case. The remaining five points contain most of the actual engineering, and they do not yield to scaling as cleanly as language did, because the physical world offers no clean training corpus and punishes every error with broken product and downtime.

Why embodiment will not simply follow the LLM curve

The tempting analogy — "VLA is doing for robots what GPT did for text, so expect the same exponential" — is half right and dangerously half wrong. What transferred: the architecture, the pretraining mindset, the generalization. What does not transfer: the data. LLMs bootstrapped on a pre-existing, permissionless, planet-scale text corpus. There is no equivalent for dexterous physical manipulation. Every hour of robot-action data is expensive to collect, and simulation — the obvious workaround — reintroduces exactly the sim-to-real gap we just discussed. This is why the field is pouring effort into egocentric human video, teleoperation datasets, and "world models": they are all attempts to manufacture the training corpus that text got for free. The bottleneck moved from algorithms to physical data, and physical data does not scale by spinning up more GPUs.

What I would actually take away

  • The paradigm shift is real; the timeline is contested. VLA is a genuine inflection. General-purpose humanoids doing arbitrary work at scale is not 2026, and probably not 2027. Bounded tasks in structured environments (warehouses, exactly) are happening now.
  • Grade every claim by metric type. Throughput in a named customer’s operation >> deployment count >> cumulative units produced >> demo video. If a company leads with the weakest of these, that choice is itself information.
  • Watch reliability curves, not capability demos. The question that predicts commercial reality is not "what can it do once" but "what fraction of attempts succeed unattended." Ask for the field number, not the lab number.
  • The data bottleneck is where the moats form. Whoever solves scalable physical-action data — not whoever has the flashiest robot — is positioned to win. Follow the datasets and the teleop fleets, not the unveilings.

Embodied AI in 2026 is neither the imminent robot workforce of the keynotes nor the perpetual vaporware of the skeptics. It is a real technology crossing a real threshold in a few narrow domains, dragging a very hard reliability problem behind it, and generating far more confident numbers than it has audited facts. The right posture is neither hype nor dismissal. It is: believe the totes, question the unit counts, and watch the second decimal place of the success rate. That is where the future actually gets decided.

Håkon Berntsen

About the Author

Håkon Berntsen is a Systems Architect at MediVox AS with over 20 years of experience in IT development, systems architecture and artificial intelligence. He is also Chairman of Open Info and an expert in AI agents and autonomous systems.