Back to recaps

Noam Brown on Why AI Benchmarks Need a Compute Budget

  • Agents
  • AI Engineering, Software, And Developer Tooling
  • AI For Science
  • AI Infrastructure, Compute, Chips, And Energy
  • Frontier Models And Capabilities
  • Policy, Governance, And Geopolitics

Watch the recap

The problem is we're in a world now where the capability of the model is a function of how much money you put into it.

The Erdős unit distance problem is an 80-year-old math puzzle proposed by Paul Erdős in 1946. It asks how many pairs of points on a flat plane can be exactly one unit apart.

Noam Brown benchmark-budget recap thumbnail
The AGI Post recap thumbnail for Noam Brown on budget-aware AI benchmarks. Source: The AGI Post YouTube

Ideas

  • No PriorsOpenAI

    Model scores need a budget axis

    Idea

    Brown argues that modern model benchmarks should use an explicit token, cost, or time budget, or plot performance against test-time compute.

    03:42 - 04:39
  • No PriorsOpenAI

    Long-running agents can outrun the model release cycle

    Idea

    Brown says the only way to know what an agent can do after a month may be to run it for a month, but new models can arrive before labs or users finish finding the old model ceiling.

    14:21 - 17:16
  • No PriorsOpenAI

    Some capabilities may already be latent but expensive to reveal

    Idea

    Brown uses the Erdos unit distance example to argue that current or recent models may have capabilities that surface only with steering, strategy scaffolds, and enough inference spend.

    17:14 - 20:41

Tags

  • Test-Time Compute
  • Benchmarks And Evaluation
  • Frontier Models
  • Research Labor Productivity
  • Agent Orchestration
  • AI Safety Governance
  • Inference Infrastructure
  • Agent Infrastructure
  • AI Policy And Governance
  • Agent Runtime
  • Software Reliability And Verification