Humanoid robots demonstrate picking and sorting abilities during the 2026 World AI Conference and High-Level Meeting on Global AI Governance in Shanghai, east China, July 19, 2026. [Photo/Xinhua]
AI's shift from the digital world to real-world applications, and from large models to AI agents, was a major theme at the 2026 World Artificial Intelligence Conference in Shanghai last month.
Zhu Konglin, a professor at the School of Artificial Intelligence at Beijing University of Posts and Telecommunications, sees this as a move from conversational AI to task-driven agents and, ultimately, "physical AI."
Unlike traditional large language models, which primarily serve as conversational partners capable of understanding and generating text, images and other content, AI agents are designed to become "actors." They can understand tasks, break them down and execute them, potentially interacting with software, hardware and, eventually, the physical world.
That transition, however, also introduces new challenges, including hallucinations, unpredictable real-world conditions and questions of explainability, accountability and risk control, Zhu said.

A self-driving vehicle navigates traffic. [File photo]
Zhu, whose research has long focused on autonomous driving, regards self-driving vehicles as a particularly demanding test of AI's ability to enter the physical world. Unlike generating text or an image, autonomous driving requires a continuous loop of perception, decision-making and action.
"The core challenge is how to enable AI to perceive its surroundings and understand the real world in a way similar to humans," Zhu said.
The physical world is open-ended and unpredictable, making it difficult for AI to make reliable judgments in rare or extreme scenarios. Speed is another issue. Autonomous driving requires near-instant decisions, but the reasoning process of current large models remains relatively slow, he noted.
Physical constraints add another layer of complexity. An AI agent controlling a vehicle cannot simply produce a theoretically reasonable decision; its actions must conform to the vehicle's dynamics and other physical limitations. Safety is equally crucial. An autonomous driving system must operate safely both when components fail and when they work as designed but encounter unexpected conditions, while keeping its actions within defined safety boundaries, he stressed.
Zhu identified three major bottlenecks: data, models and tools.
Data from rare and extreme scenarios are difficult to collect in sufficient quantities, leaving models vulnerable to unpredictable failures. Large models also remain prone to hallucinations and are often opaque, making it hard to meet explainability and reproducibility requirements for vehicle-grade systems. AI agents, along with the tools and skills they rely on, still struggle to deliver efficient real-time responses under physical constraints, while their safety boundaries remain difficult to control, he said.
The answer, he argues, lies in building trustworthy AI for autonomous driving, with trustworthy data, models and tools. This includes world models capable of predicting unfamiliar traffic situations, highly reliable edge-based driving models that balance reasoning ability with latency and determinism, and technologies for automatically identifying and simulating rare scenarios. A trusted safety architecture integrating perception, decision-making and control, backed by multiple layers of safety protection, could help AI move from simply "being able to think" toward "being able to avoid mistakes," he said.
The same principles extend beyond autonomous driving to embodied intelligence and humanoid robots. All these systems need to perceive and understand the physical world and establish a closed loop of sensing, modeling, planning and execution, he added.
These systems are likely to share common technological foundations, including world models, 3D multimodal perception and simulation-based training, as well as real-time performance, safety and robustness at the edge, Zhu noted.
In the longer term, Zhu expects a layered general-purpose intelligence system for the physical world to emerge. Cognitive foundations, basic models and simulation platforms would be shared across different applications. At the same time, motion control and scenario-specific rules would still need to be adapted to each system's physical characteristics and safety requirements.
Over the next five to 10 years, Zhu expects AI agents to gain ground first in relatively predictable, valuable and controllable environments, including autonomous freight transport, mining trucks and urban assisted driving, factory inspection and sorting, and software agents for government and corporate operations. Specialized humanoid devices could also find applications in elderly care and commercial services.
For China, he says, a key priority is to expand the supply of physical-world data, including three-dimensional, physical-interaction and real-world operational data, while improving underlying hardware and software such as reliable actuators, real-time control systems and simulation engines. Safety evaluation, usage controls and industry standards also need to keep pace.
As AI evolves from a conversationalist into an actor, technology alone will not be enough. Clear responsibility mechanisms will be needed to determine whether failures result from algorithms, hardware or improper use, and to clarify the respective responsibilities of companies, developers and users, Zhu said.
Safety testing and entry standards should be established before physical AI systems are deployed in high-risk environments, rather than after accidents occur. Ethics and law must also keep pace, particularly in areas such as human-machine boundaries, data collection and traceability, he added.


Share:


京公网安备 11010802027341号