Bryan Catanzaro On Why NVIDIA Builds Nemotron
- Agents
- AI Engineering, Software, And Developer Tooling
- AI Infrastructure, Compute, Chips, And Energy
- Capital, Markets, And Business Models
- Frontier Models And Capabilities
- Jobs, GDP, And Economic Growth
- Open Models
- Policy, Governance, And Geopolitics

NVIDIA has to deeply understand everything about how AI works. That's how we co-design all of the systems and software for our main product line.
Bryan Catanzaro started at NVIDIA in 2008, when training AI models on GPUs still sounded strange to a lot of machine-learning people. He worked on compilers and libraries for AI on the GPU, including the path that led to cuDNN. Then Andrew Ng asked him to help build Baidu's Silicon Valley AI Lab, where Catanzaro got closer to real AI applications and worked with people including Dario Amodei. Jensen Huang brought him back to NVIDIA in 2016 to build an applied research lab. The first big project became DLSS, using AI to make graphics faster and better. Around the same time, Catanzaro started Megatron because he thought text models would lead to better reasoning. Megatron was a systems project: prove that the largest Transformer models could train on NVIDIA hardware, not only on Google's TPUs. That work became part of the foundation for Nemotron.
That arc explains why NVIDIA builds open models. Catanzaro says Nemotron has two jobs. First, it helps NVIDIA learn what future AI systems need, so the company can build better GPUs, networking, compilers, software, and inference systems. Second, it keeps the open AI ecosystem strong, so companies can build their own AI close to their private data, workflows, customers, and guardrails. The logic is simple: NVIDIA sells the systems AI runs on, so it has to understand the workload from the inside.
Ideas

Nemotron Has Two Jobs For NVIDIA
IdeaCatanzaro says Nemotron helps NVIDIA understand future AI systems well enough to design GPUs, networking, compilers, software, and inference around them. It also supports an open AI ecosystem where customers and developers can build their own systems instead of depending on one model provider.
00:22:24-00:24:23
Organizations Will Run At The Limit
IdeaIn the NVFP4 discussion, Catanzaro says intelligence is valuable enough that organizations will hit a binding limit: money, servers, power, or another constraint. NVFP4 is NVIDIA's four-bit floating-point format, and his point is that once force is exhausted, more intelligence has to come from using the existing system more efficiently.
00:37:21-00:38:07
Blackwell Went All In On MoE As System Co-Design
IdeaCatanzaro says mixture-of-experts routing sends tokens through selected expert parts of a larger model, and that pattern becomes a GPU communication problem. He ties Blackwell/NVL72 to that need by describing up to 72 GPUs reading and writing each other's memory as tokens route dynamically through experts.
00:42:42-00:45:55
Coding Has A Clearer AI Training Scoreboard
IdeaCatanzaro says coding is unusually useful for AI training because it has economic value, abundant tokens, tools, and checks that can verify whether a solution works. He expects progress in other domains to require richer reinforcement-learning environments, not just the same coding setup copied everywhere.
00:58:18-01:00:16
Compute Allocation Is A Budgeted Hierarchy
IdeaAsked how NVIDIA allocates GPUs, Catanzaro says Nemotron has a compute budget, programs contain projects, projects submit requests, and NVIDIA reviews requests and budgets on a roughly two-week cycle before deciding allocations.
01:04:45-01:05:13
Research Bootstraps From Conviction To Resources
IdeaCatanzaro says research starts with conviction, then moves through small experiments, signal, resources, people, and larger bets. He uses NVFP4 as an example where a top-down strategic opportunity still needed bottom-up researchers to make it work.
01:06:55-01:10:25
Catanzaro Rejects A Sudden Singularity Frame
IdeaCatanzaro rejects a discrete singularity because intelligence is multifaceted and contextual. He contrasts math-contest skill with CEO and musician intelligence, says raw intelligence needs context and a harness, and describes AI as an external brain that may help with intelligence-limited problems.
01:13:17-01:17:48
Tags
- AI Chips
- Agent Infrastructure
- AI Infrastructure
- Open Source AI
- Frontier Models
- Open Weights
- AI Capital Allocation
- Inference Infrastructure
- Post-Training
- AI Infrastructure Efficiency
- Workflow Automation
- Frontier Lab Business Models
- Test-Time Compute
- AI Economics
- Research Labor Productivity
- Agent Runtime
- Agent Orchestration
- AI Safety Governance
- Benchmarks And Evaluation
- Compute Supply Chain
- Enterprise AI Adoption
- Human-In-The-Loop Agents
- AI Policy And Governance
- Coding Agents
- Enterprise AI Platforms
- Labor Automation
- Developer Tooling
- Economic Growth
- Software Reliability And Verification