Dwarkesh Patel on How AI Agents Could Learn on the Job
- AI Engineering, Software, And Developer Tooling
- Agents
- World Models And Robotics
- Frontier Models And Capabilities
- Jobs, GDP, And Economic Growth
Watch the deep dive
If AIs are to develop all the skills that humans have, and even skills that humans don't have, then they need to be able to learn from information revealed in unstructured, unverifiable, and ambiguous ways from scarce amounts of real-world interaction. Because in many domains, the relevant training information simply doesn't exist in any other way.
Dwarkesh Patel says frontier labs are betting on RLVR: train agents on millions of tasks with checkable answers until they become broad problem solvers. His doubt is that verifiable is not enough. A task also has to be grindable, meaning it can be replayed many times from the same starting point. Coding can work like that. Business, politics, law, markets, and operations usually cannot.
Section 1
Section 01- Dwarkesh Patel says frontier labs are betting on RLVR: train agents on millions of tasks with checkable answers until they become broad problem solvers. His doubt is that verifiable is not enough. A task also has to be grindable, meaning it can be replayed many times from the same starting point. Coding can work like that. Business, politics, law, markets, and operations usually cannot.

Tags
- Benchmarks And Evaluation
- Coding Agents
- Frontier Models
- Human-In-The-Loop Agents
- Labor Automation
- Post-Training
- Software Reliability And Verification
- Test-Time Compute
- Workflow Automation