Physical AI: Why China Believes AI Needs a Body

cs_opinion_img
"Today's AI can compose poems, generate code, and design slides with ease—yet it still cannot turn an elderly person over in bed. " Read how Chinese scientists envision a more humanistic world, powered by AI.
July 21, 2026
Hao Su (苏昊)
Hao Su is "Haoqing" Distinguished Professor at Fudan University and Inaugural Dean of the Fudan Institute of General Physical Intelligence. Prior to Fudan, he was a tenured associate professor at UC San Diego.
WAIC
Held in Shanghai, the World Artificial Intelligence Conference (WAIC) is the world's premier event in the field of artificial intelligence, dedicated to exploring the frontiers of AI technology and driving innovation in the field.
The China Academy Picks
Top picks selected by the China Academy's editorial team from Chinese media, translated and edited to provide better insights into contemporary China.
Click Register
Register
Try Premium Member
for Free with a 7-Day Trial
Click Register
Register
Try Premium Member for Free with a 7-Day Trial

Editor’s Note

The following article is translated from a keynote speech delivered by Prof. Su Hao at WAIC 2026 on July 17. In the speech, he introduced the concept of “Physical Intelligence,” describing it as the cornerstone of AI development and a potential antidote to the hallucinations that plague today’s AI models. By giving AI a “body,” Prof. Su argues, it could help address some of humanity’s most pressing challenges, including elderly care and physical labor in dangerous environments.

Keynote Speech
Physical Intelligence: From Hallucination to Reality

Distinguished leaders, distinguished guests, good afternoon. It is my great honor to share some personal perspectives at the World Artificial Intelligence Conference. The theme of this forum is “Cornerstone,” and what I wish to discuss today is precisely the most fundamental cornerstone of intelligence—physical intelligence.

Let me state my conclusion first: I am a firm optimist regarding physical intelligence. Understanding hallucination is how we bring reality closer—hence today’s title: “From Hallucination to Reality.”

1. From Large Models to Physical Intelligence

The progress of large models is there for all to see, but why do even the most intelligent models inevitably hallucinate? I believe the most fundamental reason is this: language is merely a shadow of the world. Humans first muddled through the physical world, compressing experience into language; models have always learned from this shadow, never having seen the entity that casts it. A model can fluently describe that “a cup will break when dropped on the floor,” yet it has never had the opportunity to feel the weight of a cup. Without anchors in reality, when knowledge is wrong, it has no way of knowing. You may wax eloquent on screen, yet you cannot argue gravity away.

The hallucination of large models, to a great extent, stems from having no “body.” To escape hallucination, we must cross the boundary of the “digital world,” experience things firsthand, and submit ourselves to the judgment of the “physical world”—making predictions, taking actions, being corrected by reality. This ancient process is called experimentation, and it is the key to physical intelligence.

2. What Physical Intelligence Truly Lacks: Aggregating Scattered Knowledge

So what does physical intelligence truly lack?

What we need is a model capable of aggregating all of humanity’s knowledge about the physical world. From the very outset of cognition, physical knowledge comprises at least six “rungs”, each building upon the last.

The bottom three levels belong to the objective world—regardless of whether “I” exists, they exist.

First, knowledge about objects: The world is composed of independent, persisting objects—when a ball rolls under the sofa and out of sight, it still exists. We must first recognize “what exists” before anything else becomes possible.

Second, knowledge about states: What state are these things in right now? When the ball is under the sofa, how heavy it is, whether it is soft or hard—all this belongs to this level.

Third, knowledge about dynamics: How will the world change on its own? Balls roll, water flows, objects fall when released. “Force” resides at this level—and you cannot see it with your eyes.

The top three levels come into being because of “I”—each represents the lower three levels being bound to a “subject.”

Fourth, knowledge about function: What use is this thing to me? Handles are for gripping, cups are for holding. The same chair is “sittable” for a human but not for an ant. “Things” become “tools” because of the subject.

Fifth, knowledge about goals: What state do I want the world to be in? When the ball is under the sofa, I want it back in my hand. Where the ball is, that is state; where I want it to be, that is goal. The difference between “be” and “ought to be” in goal states—two words apart, separated by the entire subject.

Sixth, knowledge about action: Knowing what needs to be done and actually being able to do it with your hands are two different things—carrying a cup of water across a room without spilling requires precision; tying shoelaces, using chopsticks require dexterity. These skills are not written in books; they reside in the hands.

From objects to actions, the higher we go, the less these things are learned by “watching”—the more they must be “done” by hand in the physical world. Every infant’s first two years are spent climbing this ladder rung by rung—Piaget called this the “sensorimotor stage”: the foundation of human intelligence is built by hand.

But models have no such childhood—they learn primarily from records left by humans. The trouble is that these six levels of knowledge are scattered across disconnected media, and the higher the level, the less it has been recorded. Internet videos are the most abundant, yet they mostly remain at the lower rungs of the ladder—appearance and motion are visible, but force and tactile feel are out of reach; textbook formulae are the most precise, describing dynamics exactly, yet only for idealized worlds; real robot data includes force feedback and operational demonstrations, reaching directly to the upper levels, yet it is extremely scarce.

Language models were fortunate: the internet already aggregated linguistic knowledge for them. Physical knowledge has no such luxury. As Polanyi said, we know far more than we can tell. Therefore, the upper half of the ladder has not yet been systematically recorded—merely browsing web pages cannot produce complete physical intelligence. Aggregation is the scientific engineering project our generation must complete: forging the breadth of video, the precision of equations, the reality of physical machines, and the subtlety of intuition into a single model, calibrating and complementing each other, completing the ladder. For physical intelligence to move from hallucination to reality, this ladder is precisely what stands in between.

3. The Imperative of Our Era, and How It Will Arrive

Why must we climb this ladder?

The answer lies neither in corpora, nor in server rooms, but in the real world.

Today’s AI can easily compose poems, generate code, and design slides— yet it cannot turn an elderly person over in bed. The value of intelligence mostly remains in the digital “bit” world, yet the most profound human needs all exist in the physical “atom” world.

Look around us: population aging is inscribed on the demographic calendar. The number of people in need of care is growing, while the number of hands available to provide care is shrinking. On the other hand, dangerous, arduous work in high-altitude, underground, and high-temperature environments also faces labor shortages. Demand is there, but we lack manpower. Physical intelligence is not humanity’s replacement, but its collaborator—it fills the gap in manpower of physical labor like tending patients and lifting loads, liberating caregivers to focus on companionship and care. Entrusting dangerous work to machines allows humans to stand behind safety lines, exercising judgment and creativity. The mission of physical intelligence is to return humanity to people.

So how will it arrive? As a technologist, here is my personal assessment. It will not be “achieved” overnight at some product launch; it will be more like electrification in the past—first lighting up factories and warehouses, then entering shops and hospitals, finally finding its way into every home. The value of physical intelligence doesn’t have to await a final culmination, but rather emerges en route.

The most dangerous stretch on this path is the reliability gap between presentations and products. To bridge this requires not fanfare, but foundational work, accumulating data—especially the half involving dynamics and interaction. It requires industry standards, supply chains, and the toughest part—social trust. Trust can only be earned bit by bit through reliability; and this trust is also the most fundamental cornerstone for physical intelligence to take root.

4. Three Predictions

Finally, I leave you with three predictions, to be tested by the future.

First, the breakthrough for general physical intelligence lies not in model architecture, but in the aggregation of knowledge. This aggregation inherently transcends the bounds of any single institution: video footage lives scattered across the internet, equations fill textbooks, haptic sensor data resides in labs, and the intuitive know-how of manipulation rests with hundreds of millions of laborers. No single entity can assemble every rung of this cognitive ladder alone. It demands collective collaboration across the entire industry and wider society: co-building datasets, co-formulating standards, and sharing infrastructure for simulation and evaluation. Once these bodies of knowledge are fully welded into a single unified model, the physical realm will witness its own internet moment — the so-called GPT moment will merely be a byproduct of this shift.

Second, the industry’s core focus will shift from crafting dazzling demos to engineering rock-solid operational reliability. In engineering, there’s a term called “several nines”: moving from 99% to 99.9% stability, every extra nine brings an exponential surge in development difficulty. The chasm between polished demos and deployable commercial products hinges entirely on those final critical nines. Generality is the destination, yet reliability is our origin. Teams willing to put in painstaking work to nail those incremental nines will travel the farthest.

Third, general physical intelligence will transform AI from a mere reader of human science into an original creator of new knowledge. Today’s large language models have digested nearly every paper humanity has published, yet they have never run a single hands-on experiment — and novel discoveries emerge precisely from experiments. Equipped with embodied hands capable of sensing, manipulating and validating the tangible world, AI will independently formulate hypotheses, conduct iterative trials, and refine its theories round the clock. The pace of breakthroughs in new materials and pharmaceutical drugs could accelerate by multiple orders of magnitude as a result.

There is no shortcut bridging digital hallucinations and physical reality. Progress hinges not on bombastic rhetoric, but on reverence for the laws of the physical world, alongside the unglamorous rung-by-rung progress of our cognitive ladder. The physical realm is intelligence’s oldest tutor, its most impartial examiner and ultimate foundational bedrock. We choose to submit our work to its judgment.

Thank you all.

Editor: Yida

References
VIEWS BY

author_image
Hao Su is "Haoqing" Distinguished Professor at Fudan University and Inaugural Dean of the Fudan Institute of General Physical Intelligence. Prior to Fudan, he was a tenured associate professor at UC San Diego.
author_image
Held in Shanghai, the World Artificial Intelligence Conference (WAIC) is the world's premier event in the field of artificial intelligence, dedicated to exploring the frontiers of AI technology and driving innovation in the field.
author_image
Top picks selected by the China Academy's editorial team from Chinese media, translated and edited to provide better insights into contemporary China.
Share This Post

Leave a Reply