
How Physical AI Learns Across Language, Video and Action — Ming-Yu Liu
How Physical AI Learns Across Language, Video and Action — Ming-Yu Liu
Machine Learning Street Talk (MLST)
番組の概要欄(原文)
<p>The car making a left turn at the start of this episode was never filmed. Cosmos 3 generated it. Ming-Yu Liu, who leads the Cosmos research at NVIDIA, explains how one model can describe a video, generate one, and produce robot actions.</p><p><br></p><p>He walks Tim through the architecture. A vision language model reasons one token at a time; its weights then initialise a bidirectional diffusion generator for video, audio and action, and a shared temporal position scheme lines up signals that run at different rates. Ming-Yu treats "world model" as a set of tools, not one definition: forward dynamics, inverse dynamics and policy, trained together under a capacity limit so that each helps the others. He also explains why plentiful first-person human video carries over to robots, which have far less data of their own, and why a Cosmos model post-trained on the DROID dataset is a good starting point for pick-and-place policies.</p><p><br></p><p>The most practical thread is testing. A neural simulator does not need accurate success rates. It only needs to rank policy A above policy B the way the real world would, so a team can narrow down which checkpoints deserve a real trial. Cosmos Dreams applies that closed-loop idea to driving and robotics, and Ming-Yu argues that humanoids around children and pets make safety matter even more than it does for cars. The conversation ends on the Super, Nano and Edge sizes (Edge targets Jetson Thor, Orin and DGX Spark) and where to find the open weights, code and data.</p><p><br></p><p>This episode is a paid partnership with NVIDIA.</p><p><br></p><p>Learn more about Cosmos: https://nvda.ws/4cJoY1S</p><p>Explore Cosmos Lab: https://research.nvidia.com/labs/cosmos-lab/cosmos3/</p><p><br></p><p>---</p><p>TIMESTAMPS:</p><p>00:00:00 A road that was never filmed</p><p>00:02:28 Inside Cosmos 3: reasoning and generator towers</p><p>00:05:02 World models: dynamics, policy and one clock</p><p>00:08:59 Learning robot skills from human video</p><p>00:11:06 Ambiguous tasks and system 2 planning</p><p>00:12:53 Neural simulators for policy verification</p><p>00:16:41 Cosmos as a starting point for robot policies</p><p>00:19:00 Cosmos Dreams and robot safety</p><p>00:22:04 Super, Nano and Edge model sizes</p><p>00:24:24 Open models, the Cosmos repo and feedback</p><p><br></p><p>---</p><p>REFERENCES:</p><p>tool:</p><p>[00:00:13] Cosmos 3 (NVIDIA Cosmos Lab project page)</p><p>https://research.nvidia.com/labs/cosmos-lab/cosmos3/</p><p>[00:18:27] NVIDIA Cosmos GitHub repository</p><p>https://github.com/NVIDIA/cosmos</p><p>[00:22:05] Cosmos3-Edge model card</p><p>https://huggingface.co/nvidia/Cosmos3-Edge</p><p>[00:22:15] Cosmos3-Super model card</p><p>https://huggingface.co/nvidia/Cosmos3-Super</p><p>[00:22:16] Cosmos3-Nano model card</p><p>https://huggingface.co/nvidia/Cosmos3-Nano</p><p>[00:22:50] NVIDIA Jetson Thor</p><p>https://www.nvidia.com/en-us/autonomous-machines/embedded-systems/jetson-thor/</p><p>[00:22:52] NVIDIA Jetson Orin</p><p>https://www.nvidia.com/en-us/autonomous-machines/embedded-systems/jetson-orin/</p><p>[00:22:53] NVIDIA DGX Spark</p><p>https://www.nvidia.com/en-us/products/workstations/dgx-spark/</p><p>[00:24:42] Cosmos 3 collection on Hugging Face</p><p>https://huggingface.co/collections/nvidia/cosmos3</p><p>other:</p><p>[00:01:07] Cosmos-Dreams closed-loop simulators (NVIDIA SIGGRAPH 2026 blog)</p><p>https://blogs.nvidia.com/blog/siggraph-news-2026/</p><p>paper:</p><p>[00:08:54] Cosmos 3: Omnimodal World Models for Physical AI</p><p>https://arxiv.org/abs/2606.02800</p><p>[00:17:43] DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset</p><p>https://arxiv.org/abs/2403.12945</p><p><br></p><p>---</p><p>RESCRIPT: https://app.rescript.info/share/e2385948cf465f0d6a2c0930150fc3ab</p>
