Podcast episode
Open models and the future of Physical AI with NVIDIA
edge-ai gpu-supply inference open-weights robotics
TL;DR
NVIDIA's Ming-Yu Liu, VP of Cosmos Lab, joins the Practical AI podcast to discuss why NVIDIA is investing in open physical AI models and the Cosmos world-model platform. The episode covers physical AI architecture constraints, world models vs. classical simulation, and the multi-agent robotics future. Useful for anyone tracking NVIDIA's AI software stack beyond GPUs, but light on hard numbers or near-term product announcements.
What was covered
- NVIDIA's open-model strategy: Liu argues open weights (plus open training frameworks and open datasets) are necessary to enable innovation that closed APIs cannot support — particularly for physical AI use cases with heterogeneous sensor configurations (different camera counts, LiDAR setups, etc.) where fine-tuning on device-specific data is essential.
- Physical AI defined: Liu frames physical AI across three verticals — automotive, factory automation, and robotics — with humanoid robots as the most anticipated form factor. He defines physical AI as AI deployed in devices that "perturb the state of the world" and produce material outcomes.
- Inference constraints on-device vs. cloud: LLMs (autoregressive transformers) are memory-bound and benefit from batching API calls from multiple users to fill GPU compute. On an embedded device, batch size is effectively one, making LLM architectures inefficient; Liu argues this will drive different architectural choices optimized for "intelligence per watt."
- World models — Cosmos 3: Liu distinguishes world simulation (input: current observation + action sequence → output: predicted future video/audio) from world understanding (input: video → output: text explanation). Cosmos 3 fuses both, supporting text, action, audio, and video as inputs and outputs in an omni-model architecture.
- Neural simulation vs. classical simulation: Classical simulators encode physics via equations and rules; neural/world-model simulation is data-driven pattern recognition over observed physical dynamics — faster iteration for policy development without requiring real vehicle fleet deployments.
- Multi-agent physical AI future: Liu sketches a manufacturing scenario where a high-level LLM agent does inventory planning, factory-floor agents manage logistics, autonomous trucks transport materials, and robot arms handle assembly — a multi-agent hierarchy with specialized roles, but acknowledges this is "decades away."
- NVIDIA's open-model ecosystem: Resources include Hugging Face (model weights), GitHub (training frameworks, cookbooks), and developer blogs. Specific model lines mentioned: Cosmos (world models), Nemo, and others.
Notable claims & predictions
- Liu: "Open model give you the innovation power so you can customize open model to better tailor for your physical use case… in the end the primary evolution will be intelligence per watt." — framing efficiency per watt as the key metric for embedded physical AI.
- Liu: "API are great, they solve problems, but they don't give you the insights… they don't allow you to tear apart and test your idea inside" — making the case that closed API access is structurally insufficient for physical AI R&D.
- Liu: "The picture I just gave — multiple agents working together completing production tasks — might be decades away." — unusually candid timeline acknowledgment for a vendor representative.
- Liu on world models as policy backbones: "The physics-based evolution [generated by the world model] has a strong correlation to the control signal you might need to use to complete certain manipulation tasks. That correlation becomes great regularization to help you build a better policy model when data is more limited."
- Liu on iteration speed: With world models replacing real-world vehicle fleet testing, teams can evaluate self-driving policy checkpoints virtually, dramatically compressing development cycles — framed as the core value proposition of Cosmos for autonomous systems.
- Liu on ecosystem necessity: "It's difficult for one company to do it alone… to achieve that beautiful future, open models are going to play a critical role." — implicitly positions NVIDIA's open-model investment as ecosystem infrastructure, not just product.
Names mentioned (from the watchlist
- NVIDIA — Cosmos Lab VP Ming-Yu Liu as primary guest; open physical AI model strategy, Cosmos 3 world model, Nemo model line, GPU/SoC hardware context.
- Hugging Face — Named as distribution platform for NVIDIA's open model weights; host Chris Benson referenced a rumored "Hugging Face acquisition" by NVIDIA (note: not confirmed in transcript, mentioned as an aside by Benson).
- Meta AI / FAIR — Implicitly referenced when Benson alludes to "certain unnamed organizations" (likely OpenAI) that were "closing" open-model efforts, noting NVIDIA stepped into the leadership gap.
- OpenAI — Alluded to indirectly by Benson as an organization that pulled back from open models; Liu references GPT historically having open-source models.
- Jensen Huang (NVIDIA) — Not mentioned by name, but NVIDIA's accelerated computing strategy attributed to company-level direction.
- Google DeepMind / Gemini, OpenAI / GPT, Anthropic / Claude — All three named by Liu as examples of LLM agents developers might use to navigate NVIDIA's open-model documentation ("whether it's Gemini, GPT, or Claude").
Why this matters for AI operators
- Physical AI as a distinct inference regime: Operators considering edge or embedded AI deployments need to rethink architecture. The autoregressive transformer + batching model that makes cloud LLMs GPU-efficient breaks down at batch size one on a robot or vehicle. Liu's "intelligence per watt" framing signals that the next wave of specialized models for physical AI will likely diverge architecturally from frontier LLMs — relevant to anyone planning inference infrastructure for robotics or autonomous systems.
- World models as synthetic data and eval infrastructure: The Cosmos use case for self-driving policy evaluation — replacing expensive real-world fleet testing with neural simulation — has direct parallels in any domain where real-world data collection is costly or dangerous. AI operators in manufacturing, logistics, or autonomous systems should evaluate whether world-model-based simulation can compress their iteration cycles similarly.
- NVIDIA's software moat deepening: By open-sourcing not just weights but training frameworks, curated datasets, and integration recipes, NVIDIA is embedding itself into the physical AI development workflow — not just the hardware supply chain. This matters for operators choosing infrastructure stacks: NVIDIA's ecosystem lock-in may increasingly come from software and model provenance, not just GPU availability.
- Multi-agent physical AI timeline calibration: Liu's candid "decades away" comment on fully autonomous multi-agent factory floors is a useful reality check against vendor hype. Near-term deployable physical AI remains narrow, sensor-specific, and heavily dependent on fine-tuned open models — operators should scope physical AI pilots accordingly rather than planning for general-purpose robot workforces.
Full analysis
NVIDIA's Cosmos Lab VP Ming-Yu Liu went on Practical AI to make the case for open physical AI models. Three things actually matter here for someone running an AI-forward company: the "intelligence per watt" argument about why robot and car models will look different from the chatbot models you use today, the "world model" pitch for training on simulated data instead of real-world fleets, and NVIDIA quietly becoming the open-model vendor just as OpenAI pulled back.
How hard is this to undo? For most readers, there's nothing to undo yet. This is a positioning conversation, not a product you buy this month. No deadline, no price change, no contract. If you build cloud software with LLMs, this changes nothing about your Tuesday. If you're anywhere near robotics, cars, or factory automation, it tells you where the stack is heading.
What's actually being decided: whether NVIDIA owns the physical AI software layer the way it owns the GPU, and whether open weights become the default for anything running on a device instead of in a data center.
The Skeptic
Liu said the quiet part twice. NVIDIA open-sources models to "understand what GPU architectures to develop next" and to grow the pool of hardware running on NVIDIA compute. This is free software that sells chips. Nothing wrong with that, but read it straight: the generosity has a meter running.
And the headline future is decades out, by his own admission. Multi-agent factory floors, humanoids doing real work, robots that follow new instructions without retraining. He called current humanoids not ready for major commercial use. When the vendor says "decades," assume he's being optimistic. There is no near-term product here you could scope a pilot around.
The world-model-replaces-real-testing pitch is the one claim that could bite. Simulated miles are not real miles until something crashes that the simulation never imagined.
The Compute Pragmatist
The useful idea in the whole hour is "intelligence per watt," and it's real. The chatbot models you rent in the cloud are efficient only because the provider stacks thousands of users' requests together to keep the GPU busy. That's called batching. On a robot or a car, there's one user: the robot. No batching. So the expensive transformer design that makes ChatGPT cheap per query makes a robot's brain wasteful and slow.
That mismatch is a genuine fork in the road. The models that run on devices will be architecturally different from the frontier models, optimized for power draw, not raw smarts. If you're planning edge or embedded AI, don't assume a shrunk-down Llama is the answer. The economics push toward something else.
The Open-Source Advocate
The structural shift worth naming: OpenAI pulled back from open weights, and NVIDIA walked into the gap. Host Chris Benson flagged it directly, and Liu didn't dodge. For anything touching hardware, open weights aren't ideology, they're a requirement. Every robot has a different camera count, different sensors. You can't fine-tune a cloud API to your specific rig. You need the actual weights on your own machine.
That's the real reason physical AI goes open while chatbots go closed. The closed-API model assumes everyone's problem looks the same. Robotics assumes everyone's hardware looks different. NVIDIA understands this better than OpenAI does, and it's building Cosmos, Nemo, and the training recipes on Hugging Face accordingly.
The Builder
Nothing ships Tuesday. Cosmos 3 is a foundation you'd fine-tune, not a product you'd call. If you're in cloud software, close the tab.
If you're in autonomous systems, the one thing to actually test is the simulation-to-reality gap. Liu's pitch is you evaluate a self-driving policy against the world model instead of a fleet of cars with safety drivers. That compresses iteration enormously if it holds. Build a harness that scores the same policy in simulation and in the real world and measure the delta before you trust a single simulated result. The whole value proposition lives or dies on how big that gap is.
Where the council splits
The Open-Source Advocate sees a clean reason physical AI goes open while chatbots stay closed: heterogeneous hardware demands local fine-tuning. The Skeptic sees the same openness as a chip sales funnel. Both are right, and that's the point. NVIDIA's open-model push is genuinely useful to builders and a moat that deepens its lock-in beyond GPU supply into the entire development workflow. Taking the free weights can leave a team more dependent on NVIDIA than before, because every fine-tuned robot brain now runs best on NVIDIA hardware at an acceptable watt budget.
The second split is on timing. The Compute Pragmatist thinks "intelligence per watt" reshapes edge inference soon, because the batching math is hard and present today. The Skeptic notes everything with a dollar attached is decades out by Liu's own words. Both hold: the architectural pressure is real now, the humanoid payoff is far away.
What it hinges on
For your company, this comes down to one question: do you deploy AI on devices, or in the cloud? Cloud builders get nothing actionable from this episode. Anyone running models on cars, robots, drones, cameras, or factory equipment should internalize that your inference economics are the inverse of the cloud's, and that the model you deploy will probably come from an open-weights ecosystem NVIDIA is working hard to own.
Prediction: NVIDIA will release at least one new open-weight Cosmos world-model variant on Hugging Face between now and NVIDIA's GTC conference in March 2027.
Confidence: High. Stated strategy, existing cadence, and a direct chip-sales incentive all point one way.
Why: NVIDIA just shipped Cosmos 3 and Liu laid out an explicit reason the company keeps releasing open weights: every open model it puts out teaches NVIDIA what chips to build next and expands the pool of hardware running on NVIDIA compute. Open-weight releases are a funnel that sells GPUs, and NVIDIA uses GTC every March as its flagship stage for exactly this kind of model and platform news. For the release to not happen, NVIDIA would have to reverse a strategy its own VP described as central and skip its biggest announcement window, which runs against every incentive in the business.
Revisit by 2027-03-31: We're right if NVIDIA publishes a new or materially updated open-weight Cosmos model on Hugging Face by the close of GTC 2027. We're wrong if no new Cosmos open-weight release appears in that window.
The real tension sits one level up. The open weights are free; the dependency they create is not. Taking NVIDIA's Cosmos models to skip building your own simulation stack is a reasonable trade for most teams, right up until NVIDIA is the only vendor whose chips run your fine-tuned robot brain at an acceptable watt budget.
Comments