Back to deep dives

Fireworks CEO Lin Qiao says the company processes 40 trillion AI tokens a day

  • AI Engineering, Software, And Developer Tooling
  • Open Models
  • AI Infrastructure, Compute, Chips, And Energy
Thumbnail for Fireworks CEO Lin Qiao says the company processes 40 trillion AI tokens a day
Image: 40 Trillion Tokens a Day (Yes, More Than OpenAI) | Lin Qiao, CEO of Fireworks

Audio deep dive

Listen to this deep dive

Spotify

IntroductionSection 01

This *Gradient Descent* podcast interview from Weights & Biases has host Lukas Biewald talking with Fireworks co-founder and CEO Lin Qiao about why she built an inference company for specialized AI models. Qiao says Fireworks processes more than 40 trillion tokens a day, with most traffic coming from customized models and deployments rather than off-the-shelf systems. They trace that demand through fine-tuning, reinforcement learning, inference economics, security, Chinese open models, and day-zero launches. They also discuss why successful AI products can become unprofitable as they scale and how Fireworks plans to compete.

Fireworks says it processes more than 40 trillion tokens a day, mostly for customized AISource13:00

And we're building our platform towards giving control of IP and giving control of cost to every single company because those company exist for a reason. That's interesting. I think a Fireworks is a company that does a really good job running open-source models. But specialized intelligence is a little bit of a different way of looking at it. Are most of your customers actually modifying the open source models before they run them? That's a really good question. Today we process more than 40 trillion tokens a day.

Qiao expects frontier models to coexist with specialized systems built on private company data. She says Fireworks handles more than 40 trillion input and output tokens a day, 95% through customized models and deployments, though providers may count tokens differently.

  • Companies keep most valuable application and enterprise data private; only a small share is public.
  • Qiao expects millions of specialized models for individual applications and use cases.
  • Available numbers suggest Fireworks' combined volume exceeds OpenAI's or Gemini's API, but their accounting may differ.
  • Fireworks sees AI spreading from coding tools into legal, finance, recruiting, marketing, sales, and support products.

Fine-tuning millions of models requires different economics, tools, and feedback loopsSource25:51

The company need to build their own eval, right? You write software, you need to write unit test and integration test to judge how good this software. It all start from there. And then once you have that, and that's what we're going to choose to hill climb. And with eval, then you start to understand, now I'm going to write how the reward look like, which is different from eval.

Frontier labs serve a few models at enormous scale; Qiao says supporting millions of customized models is a separate business. Useful reinforcement learning also needs company-specific evaluations, rewards, and domain expertise.

  • Frontier labs recover large training investments by packaging a few models as APIs.
  • Fireworks gives experts training controls while managing rollout inference; other developers use a simpler SDK.
  • A reinforcement-learning loop tests actions, grades the results, and redirects poor exploration.
  • A company needs its own evaluation system before it can design and optimize a reward.
  • Domain feedback creates hybrid roles between model research and product engineering.

Companies should turn proprietary judgment and product data into intelligence they ownSource29:58

The knowledge or the choice or the taste or the judgment to create these companies are not unified. Are not common. Are not even commonly shared. And because of that it's really hard to be captured by a general purpose model. And because of that we believe every company should own their intelligence because they are the expert carrying a taste, that judgment that unique thinking. And that should be codified into the intelligence they own and have that intelligence further power their product, make that product even better and then start to create this flying wheel.

Every company holds taste, judgment, and knowledge that a general model cannot fully capture, Qiao argues. As software gets easier to copy, proprietary customer and product data can become company-owned intelligence.

  • Companies solve different problems, so their useful knowledge is neither uniform nor widely shared.
  • Company-owned intelligence can encode that expertise and improve the product through a feedback loop.
  • Context engineering, private-data tuning, and model routing offer complementary paths to specialization.
  • Customer intent, preferences, engagement, and business logic become the harder-to-copy moat.

An AI product can find demand and still scale toward bankruptcySource37:10

Here when we talk about pricing it's actually not per token pricing because the different model their verbosity is different. Open model tend to be a little bit more verbose. Even though if you look at the pricing everything is public usually they are 10 times cheaper. But usually they're 1.5 to 2x more verbose. The cost saving is around 5 to 6x.

Product-market fit no longer guarantees a durable business, Qiao says, because inference can erase gross margins. Teams can first prove a product with a general model, then optimize each useful task for quality and cost.

  • Customers may love and pay for a product that remains too expensive to scale.
  • Large digital companies face enormous rollout costs before an AI feature's return is clear.
  • Coding agents are pushing buyers to measure value instead of maximizing token use.
  • A tenfold open-model token discount may become a five- or sixfold task saving after verbosity.
  • Once a workload is proven, teams can choose and optimize the model that fits it.

Qiao argues that open models broaden access to intelligence and cyber defenseSource41:53

Security always have two sides, the attack and defend. So the challenge of security is if it's asymmetry. If the attack has better tool and defense side, then it's really bad. If the defense has better tool than attack side, that's really good. But usually it will get the equivalent that they are on par. That I think that's kind of healthy situation. So not saying we should encourage attacker to have better tools, but they will find other ways to acquire that. So because of that nature, I feel like open model is a way to strike that balance.

Open models can keep cyber attackers and defenders on more equal footing and limit concentrated control of intelligence, Qiao says. She treats geopolitics as a separate question while asking leading American labs for stronger open releases.

  • Security is healthier when attackers do not hold a decisive advantage over defenders.
  • Accessible models let a wider community adapt systems and improve defensive work.
  • Her Hugging Face example is explicitly presented as her understanding of the incident.
  • Open data projects accelerated analytics, recommendations, and self-driving perception; Qiao wants AI to follow.
  • Stronger American open models could help businesses build intelligence from proprietary knowledge.

Fireworks prioritizes model quality over day-zero launch speedSource52:10

We hold back the launch from our side by three days. And that three days we didn't sleep at all. The reason is the release on the weights we got and the corresponding code we got has a lot of bugs. We have a lot of internal evals and it doesn't pass our threshold. We work with vLLM and SGLang, the open source community, to fix those bugs. We fix those bugs and contribute back. So they also can fix those bugs with their community. So that took us three days. And we launched three days later. But we cannot deploy a model where we know there's an issue. And that trumps everything.

Fireworks aims for day-zero support, but Qiao says quality comes first. The company delayed DeepSeek by three days after failed internal evaluations, fixed bugs with open inference projects, and contributed the fixes back.

  • Early access varies; the company once reverse-engineered Mistral code and launched before Mistral's API.
  • The company will not deploy a model with a known issue, even if the fix delays launch.
  • Customers need fast access because switching a tuned model's backbone is a major decision.
  • Those customers judge new base models with internal evaluations, not public benchmarks.
  • The platform combines training and inference; Qiao does not describe it as merely a cloud.

Fireworks' moat combines workload-specific optimization with a small, flat organizationSource58:49

We will compete. We will absolutely compete. But I think the unique part again goes back to every company exists for a reason. And the reason for us to exist is we are squarely focused on one size fits one. We squarely focus on customization. From our belief every single company is special. And we want to deliver the special intelligence for them and that reflect in special quality, special cost and speed. We'll do whatever to optimize for that, and that's what we build our platform for.

Quality-preserving infrastructure, workload-specific optimization, and a proprietary modular engine support Fireworks' claimed moat. Pragmatic open-source choices and a small, flat organization help the company move quickly.

  • Output quality comes before speed and cost optimization.
  • Its modular engine offers more than 100,000 configurations but can adopt better open tools.
  • Current LLMs help optimize kernels but have not replaced performance engineers or discovered new methods.
  • Her founder style stays authentic while she works to explain Fireworks more clearly.
  • Six technical co-founders and their intellectual honesty form the company's foundation.

Startup decisions cannot wait for perfect dataSource1:14:00

But in a startup a lot of time there's no data because we travel and pave the path. No one have traveled. If everyone's traveling that path then you shouldn't be that company. So the feedback loop of validating is important so it's okay to say this doesn't work out and we need to shut it down. But it's not okay to not make a decision because of lack of data. Not making decision is a bad decision. So we never want to be analysis paralysis and that's why we do a lot of this simulation, and try to make the best calls and then keep adjusting.

Qiao's biggest surprise was making important decisions without enough data, then testing and adjusting them quickly. Her main regret is waiting to market Fireworks in a noisy AI market.

  • Pre-mortems support candid discussion about how the company could fail.
  • Unlike Meta, a startup creating a new path has little precedent or decision data.
  • Without enough data, she makes the best available call, validates it quickly, and stops what fails.
  • Biewald advises her to keep the marketing authentic and repeat the specialized-intelligence message.

Tags

  • Open Source AI
  • Inference Infrastructure