Fireworks CEO Lin Qiao says the company processes 40 trillion AI tokens a day
- AI Engineering, Software, And Developer Tooling
- Open Models
- AI Infrastructure, Compute, Chips, And Energy

Audio deep dive
Listen to this deep dive
IntroductionSection 01
This *Gradient Descent* podcast interview from Weights & Biases has host Lukas Biewald talking with Fireworks co-founder and CEO Lin Qiao about why she built an inference company for specialized AI models. Qiao says Fireworks processes more than 40 trillion tokens a day, with most traffic coming from customized models and deployments rather than off-the-shelf systems. They trace that demand through fine-tuning, reinforcement learning, inference economics, security, Chinese open models, and day-zero launches. They also discuss why successful AI products can become unprofitable as they scale and how Fireworks plans to compete.
Fireworks says it processes more than 40 trillion tokens a day, mostly for customized AISource13:00
Qiao expects frontier models to coexist with specialized systems built on private company data. She says Fireworks handles more than 40 trillion input and output tokens a day, 95% through customized models and deployments, though providers may count tokens differently.
- Companies keep most valuable application and enterprise data private; only a small share is public.
- Qiao expects millions of specialized models for individual applications and use cases.
- Available numbers suggest Fireworks' combined volume exceeds OpenAI's or Gemini's API, but their accounting may differ.
- Fireworks sees AI spreading from coding tools into legal, finance, recruiting, marketing, sales, and support products.
Fine-tuning millions of models requires different economics, tools, and feedback loopsSource25:51
Frontier labs serve a few models at enormous scale; Qiao says supporting millions of customized models is a separate business. Useful reinforcement learning also needs company-specific evaluations, rewards, and domain expertise.
- Frontier labs recover large training investments by packaging a few models as APIs.
- Fireworks gives experts training controls while managing rollout inference; other developers use a simpler SDK.
- A reinforcement-learning loop tests actions, grades the results, and redirects poor exploration.
- A company needs its own evaluation system before it can design and optimize a reward.
- Domain feedback creates hybrid roles between model research and product engineering.
Companies should turn proprietary judgment and product data into intelligence they ownSource29:58
Every company holds taste, judgment, and knowledge that a general model cannot fully capture, Qiao argues. As software gets easier to copy, proprietary customer and product data can become company-owned intelligence.
- Companies solve different problems, so their useful knowledge is neither uniform nor widely shared.
- Company-owned intelligence can encode that expertise and improve the product through a feedback loop.
- Context engineering, private-data tuning, and model routing offer complementary paths to specialization.
- Customer intent, preferences, engagement, and business logic become the harder-to-copy moat.
An AI product can find demand and still scale toward bankruptcySource37:10
Product-market fit no longer guarantees a durable business, Qiao says, because inference can erase gross margins. Teams can first prove a product with a general model, then optimize each useful task for quality and cost.
- Customers may love and pay for a product that remains too expensive to scale.
- Large digital companies face enormous rollout costs before an AI feature's return is clear.
- Coding agents are pushing buyers to measure value instead of maximizing token use.
- A tenfold open-model token discount may become a five- or sixfold task saving after verbosity.
- Once a workload is proven, teams can choose and optimize the model that fits it.
Qiao argues that open models broaden access to intelligence and cyber defenseSource41:53
Open models can keep cyber attackers and defenders on more equal footing and limit concentrated control of intelligence, Qiao says. She treats geopolitics as a separate question while asking leading American labs for stronger open releases.
- Security is healthier when attackers do not hold a decisive advantage over defenders.
- Accessible models let a wider community adapt systems and improve defensive work.
- Her Hugging Face example is explicitly presented as her understanding of the incident.
- Open data projects accelerated analytics, recommendations, and self-driving perception; Qiao wants AI to follow.
- Stronger American open models could help businesses build intelligence from proprietary knowledge.
Fireworks prioritizes model quality over day-zero launch speedSource52:10
Fireworks aims for day-zero support, but Qiao says quality comes first. The company delayed DeepSeek by three days after failed internal evaluations, fixed bugs with open inference projects, and contributed the fixes back.
- Early access varies; the company once reverse-engineered Mistral code and launched before Mistral's API.
- The company will not deploy a model with a known issue, even if the fix delays launch.
- Customers need fast access because switching a tuned model's backbone is a major decision.
- Those customers judge new base models with internal evaluations, not public benchmarks.
- The platform combines training and inference; Qiao does not describe it as merely a cloud.
Fireworks' moat combines workload-specific optimization with a small, flat organizationSource58:49
Quality-preserving infrastructure, workload-specific optimization, and a proprietary modular engine support Fireworks' claimed moat. Pragmatic open-source choices and a small, flat organization help the company move quickly.
- Output quality comes before speed and cost optimization.
- Its modular engine offers more than 100,000 configurations but can adopt better open tools.
- Current LLMs help optimize kernels but have not replaced performance engineers or discovered new methods.
- Her founder style stays authentic while she works to explain Fireworks more clearly.
- Six technical co-founders and their intellectual honesty form the company's foundation.
Startup decisions cannot wait for perfect dataSource1:14:00
Qiao's biggest surprise was making important decisions without enough data, then testing and adjusting them quickly. Her main regret is waiting to market Fireworks in a noisy AI market.
- Pre-mortems support candid discussion about how the company could fail.
- Unlike Meta, a startup creating a new path has little precedent or decision data.
- Without enough data, she makes the best available call, validates it quickly, and stops what fails.
- Biewald advises her to keep the marketing authentic and repeat the specialized-intelligence message.
Tags
- Open Source AI
- Inference Infrastructure