What Is Physical AI? The Technology Behind Humanoid Robots Explained
Physical AI is the application of artificial intelligence to robots that sense, decide, and physically act on the world. Distinct from traditional software AI, Physical AI robots operate in physical space, learn from experience, and adapt to novel situations.
Jensen Huang, CEO of NVIDIA, introduced the term at CES 2025: "The ChatGPT moment for robotics is coming." This positions robotics at the same inflection point that large language models reached with GPT-3 and GPT-4: the convergence of foundation models, massive data, and compute sufficient for general reasoning in a domain.
Physical AI vs Software AI vs Traditional Robotics
Software AI processes information. Large language models, computer vision systems, and recommendation algorithms operate on digital data ready for processing.
Physical AI operates in the physical world through robot bodies. It must perceive unstructured environments, reason under uncertainty, and execute actions constrained by physics.
Traditional robots follow fixed programs. An industrial robot executes the same sequence of movements millions of times with perfect consistency but cannot adapt to novel situations.
Physical AI robots learn and adapt. A Physical AI robot watches demonstrations, builds a model of a task, executes it, responds in real-time to unexpected events, and improves through experience. The central distinction: generalisation.
The Three Technical Layers
Perception: Cameras, LiDAR, tactile sensors, and proprioception build a model of the environment.
World Models: Learned representations predict consequences of actions. Foundation models like NVIDIA Isaac GR00T and Google DeepMind RT-2 encode knowledge about physics and manipulation. One model adapts to many robot bodies.
Action: Motor control, trajectory planning, real-time adjustment, and force control translate decisions into precise physical movements.
The Foundation Model Breakthrough
Robotics progress was incremental for decades. In 2023-2024, foundation models transformed robotics the same way they transformed language and vision.
NVIDIA Isaac GR00T is a foundation model for robot manipulation. Instead of engineering each task from scratch, engineers fine-tune GR00T for specific tasks and hardware. One model, many robot bodies.
Google DeepMind RT-2 is a vision-language-action model trained on demonstrations. RT-2 learns from videos of humans or robots performing tasks and transfers patterns to new robot bodies.
Physical Intelligence (Pi) is a new company dedicated to building foundation models for physical AI, founded by researchers from Google, Tesla, and OpenAI.
Before foundation models: months or years to engineer each robot or task. With foundation models: weeks or days to adapt existing systems.
Who Is Building Physical AI
Every major humanoid robotics company is transitioning from scripted behavior to learned behavior.
Figure AI — Figure 02: The most advanced humanoid as of early 2026. Uses foundation models and world models for warehouse and logistics. Commercial deployments underway.
Agility Robotics — Digit: Bipedal robot for logistics. Uses learned models for locomotion and object handling. Deployed in Amazon warehouses.
Boston Dynamics — Atlas and Spot: Atlas demonstrates physical AI frontiers (parkour, adaptation). Spot is in commercial industrial inspection.
1X Technologies — NEO: Lightweight humanoid using language-based control. Early pilots with manufacturing and logistics partners.
Unitree — Unitree G1: Low-cost humanoid using foundation models. Accessible to researchers and small businesses.
Comparisons
Why Physical AI Matters
Robotics has been "10 years away" from general-purpose automation for 50 years. The bottleneck: the gap between "a robot that performs one task in a controlled environment" and "a robot that performs multiple tasks in unstructured environments."
Physical AI closes this gap because:
- Learned models are more flexible than programmed rules. Learned systems encounter variation and adapt. Programmed systems break when facing unanticipated situations.
- Foundation models reduce engineering time. Each new task takes weeks instead of months.
- Data multiplies capability. Foundation models improve as more data becomes available. New robots benefit from data collected by all previous robots.
- Transfer learning enables bootstrap. A robot trained on 1,000 bin-picking videos learns faster than a robot trained only on its own experience.
Physical AI transforms robotics from "narrow specialists in controlled environments" to "broad generalists in unstructured environments." That unlocks trillion-dollar markets.
Timeline: The Current Moment
2022-2023: GPT-2 moment — Large language models generated coherent but sometimes unreliable text. Impressive but skeptical.
Current (2026): GPT-2 equivalent for robotics — Figure 02 and Unitree G1 perform complex tasks with minimal scripting. Impressive but still requiring human intervention. Reliable only in constrained conditions. This is where physical AI stands in 2026.
2026-2027: GPT-3 equivalent — Foundation models mature. Gap between research and commercial deployment narrows. Deployments accelerate from hundreds to thousands.
2027-2029: GPT-4 equivalent — Robots handle unstructured tasks with minimal supervision. Deployment accelerates from thousands to tens of thousands. Economics shift: robots become cheaper than labour in more domains.
2030+: Post-GPT-4 moment — Robots capable of handling complex, multi-step tasks in truly unstructured environments.
FAQs
Is Physical AI the same as embodied AI? Largely yes. Embodied AI emphasises that reasoning is grounded in physical experience. Physical AI emphasises that systems operate in and directly influence the physical world.
Does Physical AI require a humanoid body? No. Physical AI applies to any robot that learns and adapts. Humanoids are the focus because the human form is general-purpose and maps to human-designed environments.
What data trains Physical AI robots? Foundation models require millions to billions of examples: human demonstrations, simulation data, and robot experience from deployed systems. The more data, the better the model.
Are Physical AI robots more expensive? Currently, yes. 2-3x more than traditional robots. As foundation models mature and volume scales, this premium narrows.
Can Physical AI robots do multiple tasks? Yes. One of the key advantages is multi-task capability. Instead of programming each task separately, robots switch between tasks learned from a single foundation model.
When will Physical AI robots be common? 2027-2029: hundreds of thousands deployed. 2029-2032: millions deployed. 2032+: mainstream adoption across manufacturing, logistics, agriculture, healthcare.
The Defining Moment
Physical AI is the correct framing for what is happening in robotics right now. The robots that matter in 2026 are not better programmed — they are learning. That distinction is what makes this moment different from every previous robotics wave that promised general-purpose robots and delivered narrow automation.
Every humanoid robot in the Geppetto directory is a physical AI platform. Browse all at the humanoid category.
The ChatGPT moment for robotics is not coming. It is already here.