Jensen Huang called 2026 the "ChatGPT moment of Physical AI". ICRA researchers in Vienna are presenting this week systems that were still science fiction three years ago. Yet no humanoid robot can reliably retrieve a glass from a dishwasher today. How far apart are promises and reality — and what still needs to happen?
Jensen Huang stood before thousands of spectators at CES in January and said 2026 would be the "ChatGPT moment of Physical AI". The phrasing was calculated. Huang knows what he's doing when he coins a term.
Physical AI is Nvidia's term for something research has been attempting for decades under various names: teaching machines not only to process the physical world, but to act within it. Not caption images. Not summarize text. But grasp, navigate, assess, react — in a world that constantly changes and doesn't wait for a clean input form.
The difference from what has been called AI so far is fundamental. A language model operates in a completely controlled environment: text in, text out, everything digital, everything predictable. A robot in a warehouse gets a camera image that might be blurry, taken in poor light, showing a box that's positioned slightly differently than the 50,000 training examples. It still has to grasp. And it has to do it now.
That's the problem Nvidia is addressing with the Isaac-GR00T platform and the Jetson-Thor chip. The idea behind it is elegant: instead of training each use case separately, a base model is pre-trained on simulation and real data — similar to how GPT-4 is trained on billions of texts — and then fine-tuned for specific tasks. The difference from old robotics development is that the model can generalize. It hasn't just learned to place Box A in Shelf B. It has learned what grasping means.
How far that reaches is shown by ICRA 2026 in Vienna, which is currently underway. Over 7,000 researchers are presenting work this week on navigation, manipulation, and what they call "loco-manipulation" — the integration of movement and grasping as a single skill rather than two separate systems. The fact that this concept even needs its own name says something about how difficult the problem is.
Autonomous navigation is the area where industry has come the furthest. SLAM — Simultaneous Localization and Mapping — has been functioning reliably in controlled environments for years. Warehouse robots have been driving through Amazon fulfillment centers since 2012, autonomous delivery robots are operating on sidewalks in dozens of cities. The challenge is no longer whether a robot can navigate a known space. The challenge is what happens when that space changes. When someone moves a chair. When a child sits on the floor. When the hallway is wet.
Foundation models for navigation — large pre-trained models, similar to LLMs but for motion planning — are the most active research field right now. NAVER Labs Europe is working on it, Unitree and Boston Dynamics are integrating them into their systems, ETH Zurich is researching in this area. The status: they work better than classical SLAM systems in unknown environments, but not yet well enough for unsupervised operation in real households.
And then there's manipulation. The actual problem.
The human hand has 27 degrees of freedom and thousands of touch receptors. It's connected to one of the largest regions of the motor cortex. When a person takes a raw egg from the refrigerator, a real-time control loop runs that integrates pressure, temperature, position, and weight in milliseconds — without us thinking about it. No humanoid robot can do that today. Most demo videos showing robot hands folding clothes or opening bottles run under very controlled conditions: perfect lighting, known objects, often slightly prepared environments.
Figure AI's Helix-02 model showed real progress in spring 2026 — Figure calls this "full-body autonomy". The robot can now coordinate its upper and lower body without both systems being controlled separately. That's not trivial. But it's also still far from doing the dishes.
What that means for household robots: the timeline the industry is communicating — first mass solutions from 2028, broad availability from 2032 — is not unrealistic, but it assumes that manipulation problems will be solved in the next three to four years. Cost won't be the issue. If Schaeffler, BMW, and Amazon take the unit volumes they've announced, actuators and sensors will become cheap enough for the consumer market. But cheap hardware is useless if the software can't do what it's supposed to.
The most realistic entry case for household robots is not the universal helper that can do everything. It's the machine that does three things very well — vacuuming, carrying, one defined transport task — and doesn't try anything else. Roomba's successor, not Rosie from the Jetsons.
For the Swiss market, that's a sober but precise perspective. If you're investing in industrial robots today or trading with them, you're in the most mature and reliable part of this market. The household revolution is coming — but it still needs a few years to catch up with physics.
This article was created with the support of artificial intelligence and editorially reviewed. The article image is an AI-generated symbolic image, not a press photo.