Dwarkesh Patel on How AI Agents Could Learn on the Job
- Agents
- AI Engineering, Software, And Developer Tooling
- Frontier Models And Capabilities
- Jobs, GDP, And Economic Growth
- World Models And Robotics
Watch the recap
If AIs are to develop all the skills that humans have, and even skills that humans don't have, then they need to be able to learn from information revealed in unstructured, unverifiable, and ambiguous ways from scarce amounts of real-world interaction. Because in many domains, the relevant training information simply doesn't exist in any other way.
Recap
Dwarkesh Patel says frontier labs are betting on RLVR: train agents on millions of tasks with checkable answers until they become broad problem solvers. His doubt is that verifiable is not enough. A task also has to be grindable, meaning it can be replayed many times from the same starting point. Coding can work like that. Business, politics, law, markets, and operations usually cannot.

Ideas
- Grindability Matters As Much As Verifiability00:02:12-00:06:10
Dwarkesh argues that a domain must be not only verifiable but also grindable: it needs replayable, parallel training targets. Coding is grindable; real-world domains like politics, business, law, markets, and operations are much harder.
- Deployment Data Has To Get Back Into The Weights00:08:41-00:15:22
Dwarkesh argues that deployment contains the most valuable learning data, but in-context learning does not scale as durable memory and gradient updates remain sample-inefficient. He presents on-policy self-distillation as one candidate way to compress session learning back into weights.
- RLVR Generalization Is The Open Empirical Bet00:06:10-00:08:41
Dwarkesh says labs are betting RLVR-trained agents will generalize from containerized tasks to open-ended real-world work, but he treats that as open rather than settled.
- The Next Paradigm Makes Deployment The Training Source00:17:23-00:19:53
Dwarkesh sketches a scenario where RLVR produces agents competent enough for real-world work, longer contexts let them accumulate week-long experience, and successful sessions are distilled back into future models.
- Dreaming Turns Scarce Experience Into Simulated Practice00:15:22-00:17:23
Dwarkesh describes a speculative training axis where a model builds simulations of a real-world task, rehearses inside them, and uses extra compute to turn scarce real-world experience into many simulated samples.
- The Lab Bet Is Millions Of Verifiable RL Tasks00:00:00-00:02:12
Dwarkesh says the frontier-lab bet is that millions of verifiable RL tasks across diverse environments can create general problem-solving agents, while skeptics may still care about sample efficiency and continual learning.
Tags
- Benchmarks And Evaluation
- Coding Agents
- Frontier Models
- Human-In-The-Loop Agents
- Labor Automation
- Post-Training
- Software Reliability And Verification
- Test-Time Compute
- Workflow Automation