Noam Brown on Why AI Benchmarks Need a Compute Budget
- Agents
- AI Engineering, Software, And Developer Tooling
- AI For Science
- AI Infrastructure, Compute, Chips, And Energy
- Frontier Models And Capabilities
- Policy, Governance, And Geopolitics
Watch the recap
The problem is we're in a world now where the capability of the model is a function of how much money you put into it.
The Erdős unit distance problem is an 80-year-old math puzzle proposed by Paul Erdős in 1946. It asks how many pairs of points on a flat plane can be exactly one unit apart.

Ideas


Model scores need a budget axis
IdeaBrown argues that modern model benchmarks should use an explicit token, cost, or time budget, or plot performance against test-time compute.
03:42 - 04:39

Long-running agents can outrun the model release cycle
IdeaBrown says the only way to know what an agent can do after a month may be to run it for a month, but new models can arrive before labs or users finish finding the old model ceiling.
14:21 - 17:16

Some capabilities may already be latent but expensive to reveal
IdeaBrown uses the Erdos unit distance example to argue that current or recent models may have capabilities that surface only with steering, strategy scaffolds, and enough inference spend.
17:14 - 20:41
Tags
- Test-Time Compute
- Benchmarks And Evaluation
- Frontier Models
- Research Labor Productivity
- Agent Orchestration
- AI Safety Governance
- Inference Infrastructure
- Agent Infrastructure
- AI Policy And Governance
- Agent Runtime
- Software Reliability And Verification