Stephen Balaban on the energy-to-tokens machine behind AI compute
- AI Infrastructure, Compute, Chips, And Energy
- Capital, Markets, And Business Models
Watch the recap
The big thing is that cloud compute is not a commodity service. It is a very complicated highly vertically integrated type of service.
Matt Turck interviewed Stephen Balaban, Lambda co-founder and CTO, in a June 18, 2026 episode of The MAD Podcast about why AI compute did not become simple GPU rental. Balaban's core claim is that an AI cloud has to join entitled land, construction, power, cooling, high-performance computing design, networking, storage, virtualization, customer-facing cloud software, financing, and customer demand into one usable service. That is why he says GPU rental indexes can mislead when they mix short-term on-demand prices with long-term contracts: the harder question is whether a provider can turn chips into reliable, partitioned, customer-ready compute.
Demand is the reason Balaban thinks the market is still generally underbuilt. If better models make more work worth doing with AI, compute demand can expand with capability. Even a 10x efficiency gain may simply let customers process 10x more tokens with the same fixed compute, where tokens are the units models process and generate. The bottleneck then becomes physical as well as technical: land, power, data-center shell, generators, UPS systems, and mechanical, electrical, and plumbing gear. Balaban says communities have real concerns about large projects, while arguing that some data-center criticism overstates water use because modern direct-to-chip liquid cooling can use dry coolers with very low evaporation.
His clearest explanation is the energy-to-tokens chain. Energy inputs become electrical power; the data center loses some power to cooling and overhead, often tracked through PUE, or Power Usage Effectiveness; servers, networking, and storage turn the remaining power into FLOPS, the raw operations used for training and inference; and model builders then care how much of that theoretical compute does useful work, a measure Balaban discusses through MFU, or Model FLOPs Utilization. At the end, FLOPS become tokens per second, and the application tries to turn those tokens into useful intelligence.
The platform layer is what makes the chip usable. Balaban says NVIDIA's moat is not only silicon but CUDA, cuDNN, networking, system reliability, and developer ecosystem gravity. Lambda's one-click cluster product is the same idea from the cloud provider side: the customer sees a simpler abstraction, while the provider has to coordinate GPU servers, CPU orchestration servers, storage, ordinary traffic networks, monitoring networks, and a compute fabric that can move data between GPUs without unnecessary CPU copies. Balaban calls that an immense software undertaking.
The finance section turns the same thesis into a balance-sheet story. Balaban says customer off-take agreements, special-purpose vehicles, private credit, and asset-backed lending can finance GPU deployments when lenders trust the customer cash flows and the NVIDIA chip assets. His sharpest example is that some H100s deployed in 2023 can lease for more later because demand stayed high and useful life looked longer than skeptics expected. He then connects that infrastructure business back to Lambda's origin story, from facial-recognition work and hardware experiments to a roughly $60,000 GPU purchase that became a cloud wedge, and forward to a future of neural software, agents, gigawatt-scale AI factories, and compute broad enough to support "one person, one GPU."
Ideas

GPU compute is not a commodity because AI cloud is vertically integrated infrastructure
IdeaBalaban says AI cloud is not commodity GPU rental; it spans land, construction, HPC design, virtualization, and cloud services.
01:48-02:27
A 2023 H100 can lease for more today because compute is becoming financeable infrastructure
IdeaBalaban says some H100s deployed in 2023 lease for more now than when first deployed because demand and useful-life assumptions changed.
40:28-42:18
The industry may still be underbuilding compute because better models expand the market
IdeaBalaban says AI compute is still generally underbuilt because scaling laws keep expanding what models can do.
07:12-09:04
Ten-times efficiency may mean ten-times more token use, not less compute demand
IdeaBalaban says a 10x model efficiency gain may let people process 10x more tokens rather than lowering total compute demand.
09:37-10:16
The real bottleneck is land, power, shell, and the public legitimacy to build
IdeaBalaban says the broad bottleneck is entitled land, utility power commitments, data-center shell, and MEP infrastructure.
10:49-11:28
NVIDIA's moat is not just chips; it is software, networking, and ecosystem gravity
IdeaBalaban says NVIDIA's advantage includes CUDA, cuDNN, networking, developer adoption, and system reliability.
26:17-28:59
One-click clusters are the customer-facing abstraction over hard infrastructure
IdeaBalaban says Lambda turns large GPU clusters, storage, networking, and orchestration into a customer-facing one-click cluster product.
28:59-34:46
AI may not just write software; software may become neural
IdeaBalaban argues that AI may not merely generate code; more software may become neural systems trained from data and feedback.
59:29-01:04:25
AI factories need off-take agreements, SPVs, and credit structures
IdeaBalaban says large AI compute projects can be financed through customer off-take agreements, special-purpose vehicles, private credit, and asset-backed lending.
38:40-40:31
Lambda's small GPU bet became the wedge into AI cloud
IdeaBalaban says a roughly $60K GPU purchase helped become the initial wedge for Lambda's cloud business.
48:59-52:00