In Physical AI Era, NVIDIA Redefines Autonomous Driving With ‘Three Computers and A Five-Layer Cake’
- Behind this strategy lies Jensen Huang's obsession with "Physical AI" – he believes that autonomous driving is the product most likely to achieve large-scale deployment within physical AI.
- NVIDIA defines intelligent driving as a “three computers” problem – in-vehicle inference computer, cloud training computer and simulation computer.
- Building upon the “three computers”, NVIDIA further refines its autonomous driving ecosystem layout into a “five-layer cake”. The core logic of this layered architecture is to use open-source models combined with diverse hardware, tools and capabilities to cover the complete ecosystem from development to delivery.
NVIDIA is positioning itself as an “ecosystem player” in the autonomous driving space through its “Three Computers and A Five-Layer Cake” strategy. This insight emerged from an exclusive interview that NVIDIA's VP of Automotive, Wu Xinzhou, gave to Counterpoint Research in April 2026, ahead of the Beijing Auto Show 2026, detailing the latest progress and strategic layout of the company’s intelligent driving business.
Behind this “Three Computers and A Five-Layer Cake” strategy lies Jensen Huang's obsession with "Physical AI" – he believes that generative AI, agentic AI and physical AI will drive the Fourth Industrial Revolution, and autonomous driving is the product most likely to achieve large-scale deployment within physical AI. Considering that the vehicles across the globe travel around 13 trillion miles annually, if every mile could be automated in the future, the commercial value would be immeasurable.
1. Technical Architecture: Full-Stack Layout of ‘Three Computers and A Five-Layer Cake’
NVIDIA defines intelligent driving as a “three computers” problem:
- In-vehicle inference computer: Responsible for real-time perception and decision-making.
- Cloud training computer: Responsible for model training and data annotation.
- Simulation computer: Responsible for large-scale scenario validation.
These three computers form a closed loop, spanning in-vehicle data collection, cloud model training, simulation validation and deployment back to vehicles. NVIDIA has capabilities across all three segments, thus building a complete toolchain.
Building upon the “three computers”, NVIDIA further refines its autonomous driving ecosystem layout into a “five-layer cake”:

The core logic of this layered architecture is to use open-source models combined with diverse hardware, tools and capabilities to cover the complete ecosystem from development to delivery, enabling automaker customers to benefit regardless of which segment they participate in.

2. Core Technology Breakthrough: The Paradigm Revolution of Reasoning-Based VLA Model
One of the key selling points of NVIDIA's Alpamayo model is its "reasoning-based VLA model”. The dilemma of traditional end-to-end solutions lies in the need for massive amounts of driving data to cover various corner cases, resulting in high costs and low efficiency in achieving autonomous driving. The reasoning-based VLA model introduces two key capabilities to address the pain points of traditional solutions:
- Foundation Model: Alpamayo's backbone is distilled from the Cosmos world model, which is trained on internet-scale video data, naturally possessing an understanding of the physical world. Although the foundation model only has 80,000 hours of driving data, its "parent" has learned massive physical laws, giving it far superior generalization capabilities than models of similar scale.
- Language Reasoning Capability: Humans can understand through text “there’s an animal ahead, don't hit it”, without having seen one in person. Similarly, after introducing reasoning capabilities, the model can handle unseen scenarios through the language's generalization ability, significantly reducing data requirements.
Wu Xinzhou used a vivid analogy to illustrate this: "A novice learning to drive only needs to read a driver's manual and drive for 20-30 hours to get a license, because they already have an understanding of the physical world (foundation model) and language generalization ability (reasoning ability). Future autonomous driving models will achieve similar capabilities, possibly requiring only dozens of hours of training data."
In the end-to-end era, traditional modular validation methods become ineffective. NVIDIA's solution is "neural reconstruction (Omniverse NuRec)” – reconstructing pixel-level scenarios of the physical world through models, editing new elements within them (such as scooters, traffic cones), and then changing weather, lighting and time through Cosmos, achieving multiplied scenario validation efficiency.
The core values of this toolchain are:
- Two million simulation validations daily, far exceeding the efficiency of real vehicle testing.
- A five- to tenfold improvement in data validation capability, significantly accelerating model iteration.
- Production-grade quality, already available to ecosystem partners.
3. Commercialization Progress: Dual-Track Approach of L2++ Production and L4 Vision
In Counterpoint's view, although NVIDIA's L3/L4 layout is robust, L2++ remains the main source of its short-term revenue. According to NVIDIA, its collaboration with Mercedes-Benz is the benchmark case for its L2++ production, which includes the following milestones:
- 2025: Completed first-round delivery in Europe and the Americas.
- 2026: Point-to-point L2++ deployment across the US while expanding to multiple European cities.
- Second half of 2027: Achieving globalized L2++ production target.
This solution adopts a “hybrid end-to-end” architecture, where the end-to-end model handles "driveability" (imitation learning, experience close to human), while running a rule-based classic algorithm in parallel as a safety fallback.
The commercialization milestones for L3/L4 mainly appear in 2027-2028:

According to NVIDIA, the core architectural features of the L3/L4 system include:
- Dual Thor ECU redundant design: Fail-operational capability even under single-point failure
- Sensor redundancy: 14 cameras + 9 millimeter-wave radars + 1 LiDAR
- Dual algorithm parallelism: End-to-end + rule engine, ensuring both safety and experience
Wu Xinzhou also provided clear technical judgment on L3/L4: “From a technical perspective, there are no major bottlenecks. Large models and reasoning technology have already addressed past major difficulties; the current challenges are mainly at the systems engineering level and specific workload issues.” Regarding the recent industry debate on "whether to skip L3," he believes that L3 and L4 have similar technical difficulty but significantly different business models, such as:
- L3 suits individual consumers: The 10-second takeover requirement can already free drivers (such as freeing attention during highway commuting, giving time back to users themselves).
- L4 suits operational services: Due to the need for remote operation capabilities, it is unrealistic for automakers to deploy remote takeover systems across millions of vehicles in a short time.
4. Market Strategy: Ecosystem Logic of ‘Revenue from Every Mile’
NVIDIA builds a deep technological moat through "Three Computers and A Five-layer Cake”, forming a closed loop spanning cloud training, simulation and open-source models. Therefore, even if automakers develop their own chips, they still need NVIDIA's GPUs for model training and simulation toolchains to generalize model reasoning capabilities.
Wu Xinzhou explicitly stated: “We hope all future driving miles will be automated, and we participate in every single mile.” In Counterpoint's view, “revenue from every mile” is both a powerful declaration and NVIDIA's business model innovation in the autonomous driving space, comprehensively empowering automakers through a complete autonomous driving ecosystem. Regardless of which segment automakers choose (i.e., any layer of the five-layer cake), NVIDIA can gain valuable returns.
5. Future Competitive Focus: Computing Power Just a Starting Point, Ecosystem is the Endgame
Over the past few years, the core metric for autonomous driving chip competition has been TOPS computing power. However, Wu Xinzhou pointed out that future competitive focus will shift to multiple dimensions, including higher sensor data utilization, higher frame rates, reinforcement of long-term reasoning, and higher energy efficiency ratios – these are also key areas for NVIDIA's next-generation chips.
When asked about competitors, Wu Xinzhou's answer was thought-provoking: "We are still an ecosystem player. There is no real competitor; we want everyone to succeed together."
In Counterpoint's view, this is not PR rhetoric but a result of its business model – NVIDIA's revenue comes from the overall adoption of AI in the automotive industry, not single-chip market share. For example, Tesla develops its own chips but still uses NVIDIA for cloud training, and Waymo, Uber and others are all part of its autonomous driving ecosystem.
In the Chinese market, NVIDIA’s ecosystem is also continuing to grow. During Auto China 2026, NVIDIA, together with multiple automakers and NVIDIA DRIVE ecosystem partners, showcased the latest breakthroughs in intelligent mobility. The latest collaboration updates include:
- Physical AI
NVIDIA and Chery Automobile are collaborating across three domains – assisted driving, in-cabin AI and robotics – to jointly develop and deploy physical AI.
- Scaled Deployment of Autonomous Driving
Desay SV, WeRide, Pony.ai and DeepRoute.ai announced their latest progress based on the NVIDIA DRIVE Hyperion platform, accelerating the scaled commercialization of robotaxi services.
- AI Agents Powering Next-Gen In-Vehicle AI Experiences
- MediaTek's Dimensity Auto Cockpit Platform C-X1 integrates the NVIDIA Blackwell GPU architecture and deep learning accelerators, with support for the NVIDIA CUDA ecosystem.
- Alibaba's Tongyi Large Model Business Unit (Tongyi Laboratory) is running Qwen-Omni, its full-modal large model, on the NVIDIA DRIVE platform.
- Lenovo Vehicle Computing has launched the Auto AI Box, a cockpit-and-driving integrated computing platform powered by NVIDIA DRIVE AGX Thor.
- Volcano Engine is developing customized AI software solutions based on the NVIDIA DRIVE AGX Thor and Orin platforms.
- Pateo announced to leverage the AI compute of NVIDIA DRIVE AGX Thor to power its AI BOX.
- ThunderSoft announced AquaClaw, an in-vehicle AI agent operating system built on the NVIDIA NemoClaw reference software stack.
Analyst Take: ‘ChatGPT Moment’ for Autonomous Driving Has Arrived
Wu Xinzhou emphasized multiple times during the interview: "With the emergence of reasoning-based large models, the ChatGPT moment for autonomous driving has arrived." Behind this judgment is confidence in the technological inflection point. The corner case problems that have troubled the industry are being systematically solved through the combination of foundation models and reasoning capabilities. From an engineering perspective, autonomous driving is more of a "systems engineering" challenge than a “scientific problem”.
As Jensen Huang stated, physical AI will drive a hundred-fold GDP growth, and autonomous driving is the best entry point for physical AI. In this marathon of physical AI, NVIDIA has chosen to be an "ecosystem player", not pursuing a monopoly but pursuing the ability to contribute to every mile of autonomous driving. This is a longer-term business vision and a more resilient competitive strategy.
Category
Industry
Automotive, Semiconductors
Service
Standard
Report Type
Report
Time Period
Other
Receive our insightful weekly newsletter and stay ahead of the competition.
Author
Kevin Li
Kevin is an Associate Director at Counterpoint Research based in Beijing. At Counterpoint, he leads the China automotive market research. Kevin has 12 years of experience in 5G/V2X, connected vehicles, intelligent cockpits, and intelligent driving in market analysis firms, including Strategy Analytics and TechInsights. Previously, Kevin has worked for China Unicom/China Netcom as a Senior Engineer and International Cooperation Coordinator for 10 years. Kevin holds an MSc in Mobile Communications from Beijing University of Posts and Telecommunications.