
The GPU Myth: State of AI Compute 2026 | Stephen Balaban
The GPU Myth: State of AI Compute 2026 | Stephen Balaban
The MAD Podcast: How AI Gets Built — with Matt Turck
番組の概要欄(原文)
<p>Many people said GPU compute would become a commodity. The opposite happened — and a new category of "neoclouds" is now racing to build the physical backbone of the AI boom. Stephen Balaban, co-founder and CTO of Lambda, explains why the conventional wisdom was exactly wrong, why we're still massively underbuilding compute, and what it actually takes to stand up a gigawatt-scale AI factory: land, power, cooling, networking, and a financing stack most people have never heard of. We go deep on the physics of how energy becomes tokens, NVIDIA's real moat, why a 2023 GPU can lease for more today than the day it shipped, and Stephen's provocative vision of "neural software." Plus the wild Lambda origin story — from a facial recognition startup to a camera in a baseball cap to a near-billion-dollar cloud business. This is the state of AI compute in 2026, from inside one of the companies building it.</p><p><br /></p><p>(00:00) — Cold open</p><p>(01:21) — Why GPU compute was never a commodity</p><p>(02:45) — The H100 price index and what it gets wrong</p><p>(04:02) — The real moat: technology or financing?</p><p>(05:57) — Winner-take-all, or room for many neoclouds?</p><p>(06:48) — Are we overbuilding or underbuilding AI compute?</p><p>(09:26) — What if AI gets 10x more compute-efficient?</p><p>(10:44) — The real bottleneck: land, power, and shell</p><p>(11:38) — The backlash against data centers — and the misinformation</p><p>(15:00) — Opening the hood: from photons to tokens</p><p>(17:11) — Extracting more value from the same chip</p><p>(19:26) — Frontier inference and distributed training, explained</p><p>(23:26) — What actually drives compute cost</p><p>(25:21) — Lambda's chip stack and the NVIDIA relationship</p><p>(26:17) — A multi-silicon world? CUDA, CUDNN, and NVIDIA's real moat</p><p>(28:59) — Networking, storage, and the one-click cluster</p><p>(34:46) — Renting vs. owning, and full vertical integration</p><p>(36:24) — How global is Lambda? Does location still matter?</p><p>(38:44) — The financing stack: off-take agreements, SPVs, and credit</p><p>(41:16) — Why a 2023 GPU leases for more today</p><p>(42:36) — A futures market for compute?</p><p>(43:54) — Origin story: facial recognition, Perceptio, and Apple</p><p>(47:03) — The Lambda hat and Dream Scope</p><p>(48:59) — The $60K bet that became a cloud business</p><p>(52:00) — Holding the team together through the hard times</p><p>(54:30) — Bringing on a new CEO; Stephen as CTO</p><p>(57:33) — Matching xAI on high-velocity deployment</p><p>(59:29) — "AI won't write software — it will become the software"</p><p>(01:01:30) — Neural software vs. vibe coding</p><p>(01:04:25) — Do agents change the compute layer?</p><p>(01:06:14) — Self-assembling software inside Lambda</p><p>(01:08:18) — Gigawatt-scale AI factories</p><p>(01:08:57) — One person, one GPU</p><p>(01:12:04) — Hot takes: overrated and underrated in AI</p>