Deep Dives

Source-ordered deep dives into important AI research and conversations.

Source thumbnail for Baseten on the inference frontier: 200,000-token routing, 4–6× speedups, and self-optimizing AI
deep diveListen

Baseten on the inference frontier: 200,000-token routing, 4–6× speedups, and self-optimizing AI

This Latent Space podcast interview has host swyx asking Baseten’s Philip Kiely and Ali Taha how inference turns an open model into a product: request routing, quantization, GPU kernels, model parallelism, and AI video. It gets properly weird near the end, when GLM-5.2 helps rewrite its own serving code and continual learning starts to blur the line between training and inference.

  • AI Infrastructure, Compute, Chips, And Energy
  • AI Engineering, Software, And Developer Tooling
  • Open Models
The Inference Frontier: 10x Faster Models to Self-Optimizing AI — Philip Kiely & Ali Taha, Baseten