ENG
Insight

Humanoid Robots, Autonomous Vehicles Pushing AI into Physical World

0
October 28, 2025
  • A central ambition of artificial intelligence is to create agents that can interact with the physical world, a goal now being significantly advanced by the fast development of fundamental large models.
  • Humanoid robots and autonomous vehicles, as the leading edge of physical AI commercialization, are converging on a shared hardware and software paradigm that will accelerate the development of embodied AI models and edge computing, paving the way toward genuine physical AGI (artificial general intelligence).
  • Counterpoint forecasts that the compound annual growth rate (CAGR) of global humanoid robot shipments will be 69.7% from 2025 to 2030, driven by expanding use cases, declining hardware costs and advancing AI capabilities.


For decades, AI’s grand challenge has been to break free from the cyber realm and interact with the physical world. The goal of moving beyond text prompts and image generation to truly operate within our reality is now being significantly advanced by the advent of physical AI, which integrates and processes data from different sources, making it possible for machines to understand complex real-world environments more comprehensively and make corresponding decisions.
The demands of interacting with the physical world are further driving the emergence of Vision-Language-Action (VLA) models, which integrate the three core components of machine autonomy — perception (Vision), reasoning (Language) and execution (Action) — into a single large model by outputting tokenized action instructions.



Source: Paper titled ‘Vision-Language-Action Models: Concepts, Progress, Applications and Challenges’

The emergence of VLA models has brought revolutionary progress to humanoid robots and autonomous driving solutions. Traditional algorithms often rely on preset strategies or hand-engineered rules, and struggle to adapt to sophisticated and ever-changing real-world environments. By unifying the semantic space of vision, language and action, VLA models enable robots to understand scenarios, plan actions and execute tasks in a way that is closer to humans. For example, in robotics, VLA models allow machines to complete multi-step operations through natural language instructions, achieving the leap from “seeing” to “doing”. In autonomous driving, they enable vehicles to combine visual perception with semantic understanding, dynamically adjust decisions, and respond to unexpected situations.

With the integration of powerful edge computing and soaring training data, VLA models are driving humanoid robots and autonomous vehicles toward a higher level of autonomy, laying the foundation for the realization of general embodied intelligence.

Robots, autonomous vehicles underpinned by similar tech foundations

With the support of VLA models, humanoid robots and autonomous vehicles share the same perception-reasoning-execution paradigm, which further enables them to share similar hardware and software technology stacks. Both demand 3D perception and reasoning hardware – sensor fusion systems (cameras, LiDAR, radar) to enable multimodal environmental awareness, and high-performance edge computing platforms for real-time inference.

Software-wise, apart from VLA models, both also require massive amounts of data to train models built on an end-to-end architecture, and this has thus spurred the advancement of ‘world models’, used for generating synthetic data and conducting simulation testing.


The differences between humanoid robots and autonomous vehicles first lie in their functional requirements. Besides cameras, LiDAR and radar, robots also require more diverse sensors such as tactile, torque and temperature. In terms of execution units, cars can be regarded as “four-wheeled robots” with only two degrees of freedom, while robots usually have dozens of degrees of freedom and a more complex structure. With the physical forms of robots yet to be unified, different “limb” structures lead to differences in hardware and software. This high complexity also results in a scarcity of training data, making robots more dependent on the development of world models than autonomous vehicles for synthetic data.


Commercialization of autonomous devices further promotes physical AI ecosystem

The commercialization of autonomous devices is acting as a powerful catalyst, significantly accelerating the maturation of the physical AI ecosystem. A pivotal shift began around 2024, when the end-to-end autonomous driving paradigm transitioned from prototypes to practical deployment, making L2+ autonomy widely available in consumer vehicles. Since then, the paradigm rapidly advanced from end-to-end neural network approaches to incorporating Vision-Language Models (VLM), and further to the more integrated Vision-Language-Action (VLA) frameworks. Building on this momentum, the year 2025 has become a landmark for humanoid robotics. Several leading manufacturers have unveiled remarkable breakthroughs, demonstrating robots that can perform complex tasks with a notable degree of autonomous execution, moving beyond pre-programmed routines. Counterpoint forecasts that global humanoid robot shipments will reach 256,000 by 2030, driven by expanding use cases, declining hardware costs and advancing AI capabilities.


This progress in the two domains underscores a critical reality – commercialization is not merely an outcome but the essential fuel for continuous innovation in physical AI. The research and development investments required to achieve these technological leaps are immense. Consequently, large-scale product deployment and market adoption are imperative to generate the necessary revenue to amortize these costs. This commercial success creates a virtuous cycle, funding further R&D, attracting talent and driving down component prices, thereby solidifying the foundation for the entire physical AI ecosystem and propelling it toward a more intelligent and automated future.


Receive our insightful weekly newsletter and stay ahead of the competition.

Author

Shaochen Wang

linkedin_icon

Shaochen Wang is a Research Analyst at Counterpoint Research, based in Beijing, China. At Counterpoint, he closely monitors the automotive and intelligent robotics industries, with a particular focus on pivotal technologies—including Autonomous Vehicles, Humanoid Robots, and Artificial Intelligence—that underpin generalized autonomy. He began his career at Lenovo Group in the Corporate Strategy division. Prior to joining Counterpoint Research, his most recent role was as a Consultant at EY China. He holds a MSc in Operations Research & Industrial Engineering from the University of Texas at Austin, and a BSc in Civil & Environmental Engineering from Technion-Israel Institute of Technology.